trips/spatial/README.md
Greg Pomerantz df5fed691b spatial: remove the pgvector path entirely
The targeted /search + /kinds approach is the only POI search now;
nothing consumes embeddings. Drop:
- poi_vec data + the 'vector' extension (DB)
- embed/ (embedpoi), 20-vector.sql, db/ (the pgvector image workaround
  — compose goes back to plain postgis/postgis:16-3.4)
- /semantic endpoint + embed client code from spatiald
- EMBED_* env vars and README sections
- .gitignore: ignore the spatial/ go build artifact
2026-09-10 15:07:32 -04:00

118 lines
5.5 KiB
Markdown

# spatial — PostGIS source of truth
Native spatial queries for the trip planner (DESIGN.md §Data: "PostGIS
(Postgres) — import PBF via osm2pgsql; gives SQL spatial queries (radius,
nearest, bbox)"). Replaces the flat-file/Overpass approximations.
## Stack
| Piece | What it is |
|---|---|
| `docker-compose.yml` | `postgis/postgis:16-3.4` (DB, port 5432) + one-shot `iboates/osm2pgsql` importer + `spatiald` (query service, port 5005) |
| `schema.sql` | runs on first boot: `poi` table (named POIs as points, GIST index), `spatial_extract` bookkeeping, `refresh_poi()` |
| `import.sh` | loads a PBF from `../osm/` via osm2pgsql, refreshes `poi` |
| `main.go` (spatiald) | JSON query service over `poi` (radius / KNN / corridor / targeted tag+name search) |
| `Dockerfile` | builds spatiald |
## Prerequisites
* Docker with the current user in the `docker` group (on this box:
one-time `sudo usermod -aG docker $USER`, then re-login).
* PBF extracts in `~/trips/osm` (already present: `nh.osm.pbf`,
`colombia.osm.pbf`, US state extracts; override with `OSM_DIR`).
## Usage
```bash
docker-compose up -d postgis # first: runs schema.sql
./import.sh ~/trips/osm/colombia.osm.pbf colombia
./import.sh ~/osm-build/data/northeast.osm.pbf northeast # merged NE extract
docker-compose up -d spatiald # query service on :5005
```
The mock server (`mock/server.js`) proxies it: `GET /spatial/*``:5005`,
plus `GET /spatial-status` for the UI badge. When the service is up, the
chat assistant gains the `poi_near`, `poi_search` and `poi_kinds` tools
automatically.
## spatiald endpoints
All return `{"count":N,"results":[{extract,osm_id,name,kind,opening_hours,
fee,website,addr_city,dist_m},…]}`. Common filters:
* `extract=nh` — one loaded extract
* `kind=amenity=restaurant` — exact kind, or `kind=restaurant` (any family)
* `name=café` — ILIKE substring
* `limit=20` (max 200)
| Endpoint | Query | SQL core |
|---|---|---|
| `GET /health` | — | extract bookkeeping |
| `GET /near` | `lat,lng,r(m; default 500)` | `ST_DWithin(geography, …, r)` + KNN ordering |
| `GET /nearest` | `lat,lng` | `ORDER BY geom <-> point` |
| `GET /corridor` | `points=lng,lat;…`, `r` | `ST_DWithin(geom::geography, ST_Buffer(line::geography, r))` |
| `GET /search` | `kinds=a\|b`, `terms=x\|y`, optional `lat,lng,r` | indexed `kind` + FTS/trigram `name` match |
| `GET /kinds` | `q?`, `extract?`, `limit` | `GROUP BY kind` — the tag vocabulary |
`/search` and `/kinds` are the agent's main POI lookups (see the next section).
All spatial results carry `lat`/`lng` so the UI can drop map pins.
Examples:
```bash
curl 'localhost:5005/near?lat=42.35&lng=-71.06&r=1000&kind=restaurant'
curl 'localhost:5005/nearest?lat=10.40&lng=-75.54&kind=tourism&limit=5'
curl 'localhost:5005/corridor?r=300&kind=fuel&points=-71.06,42.35;-71.07,42.36'
curl 'localhost:5005/kinds?q=wine'
curl 'localhost:5005/search?kinds=amenity=bar&terms=wine%7Cvino%7Ccava&lat=4.65&lng=-74.08&r=8000'
```
## Targeted POI search (the default — no embeddings)
The OSM `kind` column is a **closed, standardized vocabulary** (~1.3k distinct
tags: `amenity=restaurant`, `tourism=museum`, `historic=fort`, …), and `name` is
a short free-form string. That's enough structure to skip vector embeddings
entirely:
* **The LLM agent is the semantic layer.** It translates the user's concept into
OSM `kind`s + local-language `name` keywords ("fortress with a view" →
`historic=fort` / `tourism=viewpoint` + `view|panoramic|mirador`). It is
already resident in VRAM for chat, so this costs nothing extra.
* **`/kinds`** returns the tags that actually exist in the extract (`GROUP BY
kind` with counts), so the agent grounds its choices in real data instead of
guessing tags that aren't there.
* **`/search`** retrieves with indexed SQL: `kind IN/ILIKE` (the `poi_kind`
index), `name` full-text (`poi_name_fts`, `to_tsvector('simple')`) and trigram
(`poi_name_trgm_ops`, `pg_trgm`) for fuzzy/substring matches, plus an optional
`ST_DWithin` radius. Ranked by trigram similarity, then distance.
This replaces the old pgvector approach for the agent: no embedding model, no
extra VRAM, no ~4.6 GB vector table, no one-time 450k-row embed job — and it is
arguably more accurate here, because `kind` is ground truth.
## Data notes
* Built for **osm2pgsql 2.x** (`iboates/osm2pgsql`): tables are
`planet_osm_point` / `planet_osm_polygon` (not the 1.x `_node`/`_way`),
geometry lives in a `way` column in EPSG:3857, and the import runs with
`-k` so tags without a fixed column (`opening_hours`, `fee`, `website`,
`addr:city`) land in the hstore `tags` column that `refresh_poi()` reads.
`refresh_poi()` also dedupes relation-derived polygons (2.x repeats them
per relation with negated ids).
* `poi` covers every **named** point/polygon with an `amenity`/`tourism`/
`shop`/`leisure`/`historic`/`place` tag. Unnamed amenities (a nameless
kiosk) are out of scope for v1.
* `kind` filter semantics: `kind=amenity=restaurant` (exact),
`kind=restaurant` (tag value, any family), `kind=tourism` (whole family).
* Corridor `points` must URL-encode the semicolons (`%3B`) — Go's
`url.Parse` drops the tail of a value that contains a raw `;`.
* Multiple extracts coexist in one DB, disambiguated by `poi.extract`
(set automatically by `import.sh`).
## Roadmap hooks
* `bbox` queries: trivial addition (`ST_Contains(ST_MakeEnvelope,…)`).
* Routing-graph join: OSRM/GraphHopper geometries can be loaded into the
same DB (`route_geom` table) for true along-route analytics instead of
the current buffer-over-polyline.