trips/DESIGN.md
Greg Pomerantz efde2cc71b maps project: design, survey, mock app, and route-aware planning backend
- DESIGN.md: full design incl. driving-trip requirements (R1-R4),
  stays model, focus mode, mobile, provenance rules
- SURVEY.md: open-source landscape
- mock/: interaction mock (Florence itinerary, focus mode, stays,
  region stops, mobile layout)
- router/: Go module (stdlib-only) with Router interface
  (Valhalla + OSRM backends), stop_cost, optimize_stops, corridor,
  routectl CLI, bench (5 real NE-corridor tasks, 26 checks passing),
  integration tests, and setup-osrm.sh for the self-hosted router
- osm/: NH+MA+CT+NY PBFs (gitignored) + setup artifacts
2026-09-06 00:05:17 -04:00

1045 lines
60 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# LLM Map Assistant — High-Level Design
An itinerary building & refinement tool: the user chats with the system to
build and iteratively refine a multi-stop trip (sightseeing, food, errands,
travel between places), and the system answers map questions — routes and
travel times, businesses, opening hours, products & services — against
authoritative geospatial data. The itinerary is a first-class, versioned
state object that the LLM edits; free-form map QA is a sub-feature.
The system carries a trip through the whole planning lifecycle: **idea
generation & filtering → organization & planning (incl. multi-modal route
planning) → changes & refinement → identification & selection of
alternatives** (see Planning lifecycle).
Deployment: **web-first** — serious planning happens on a desktop browser
(map + timeline + chat, lots of screen). The mobile app is an **on-trip**
companion: read the plan, navigate, check "open now?", make small edits —
and it must keep working in **areas with limited network coverage** (see
Interface & offline strategy). The LLM agent + tools + data live on a
server; on-device inference is a later fallback, not a target (see Model
selection).
## Core principle (from the research)
Don't let the LLM "look at" the map. Have the LLM **orchestrate deterministic
geospatial tools** and let the tools produce exact facts (distances, times,
coordinates, hours). The LLM contributes:
- understanding and decomposing the user's question,
- choosing/sequencing tool calls (agent loop),
- filtering/ranking retrieved results,
- composing a natural-language answer.
Evidence:
- *MapQA* (arXiv 2503.07871): open-domain geospatial QA benchmark built on
OpenStreetMap; question types include routing, POI attributes, spatial
relations. A good evaluation set for us.
- FOS4G 2026 talk "A Systematic Comparison of RAG Architectures for
Geographic POI Question Answering": structured/geotemporal RAG with tool
access beats naive vector RAG for spatial POI queries.
- *MapBench* (arXiv 2503.14607) and *MapReason-OSM*: even strong VLMs fail
at reliable path-finding/symbol reading on map images. ⇒ map images are
for the user, not the engine.
## Itinerary: the core state object
The product is *itinerary building & refinement*, so the itinerary is a
typed, persisted document — not chat history. It is **hierarchical**:
the top level spans countries/regions; each *place* (a country, city,
neighbourhood — **free-form, not a fixed category list**) carries its own
refined objective/"vibe" and contains its own days and stops.
```jsonc
{
"id": "trip-2026-09",
"version": 12, // every LLM edit = new version (undo, diff)
"preferences": { // the evolving taste model (see Exploration loop)
"likes": ["slow_mornings", "street_food", "coastal_walks"],
"rules_out": ["theme_parks", "early_flights"],
"budget": "moderate",
"vibe_by_place": {"portofino": "relaxed, food-first",
"florence": "art, but not the big 4 museums"}
},
"dates": ["2026-09-12", "2026-09-20"],
"modes": {"default": "foot", "car": false, "transit": true},
"places": [{
"id": "PL1",
"name": "Tuscany — Florence & coast", // free-form label, any granularity
"centroid": [43.77, 11.25], // its geo anchor (city or area)
"bbox": [43.1, 10.0, 43.9, 11.9], // for tiles/caching/planning
"vibe": "art + food, relaxed pace", // refined objective for this place
"dates": ["2026-09-12", "2026-09-17"],
"days": [{
"date": "2026-09-12",
// no per-day `base`: the anchor is RESOLVED from `stays` (below) — the
// stay(s) covering this date. >1 candidate stay on the same date ⇒
// worst-case anchoring for that date (see rules).
"stops": [
{"poi_ref": "P456", "arrive": "09:30", "depart": "10:15",
"purpose": "visit",
"suggested_duration": {"value": "02:30", "source": "search"},
"planned_duration": {"value": "01:00", "source": "user"}, // sticky override
"price": {"amount": 14, "currency": "EUR", "basis": "per person",
"source": "search"},
"slot": "morning_visit",
"kind": "point|region|transit", // region = vague ("lunch in Oltrarno")
"state": "planned|maybe|alt"}, // one `planned` per (day, slot)
{"poi_ref": "P789", "arrive": "12:00", "depart": "13:30",
"slot": "lunch", "kind": "point", "state": "planned",
"purpose": "meal", "meal": {"kind": "lunch",
"typical_duration": {"value": "01:15", "source": "search"},
"typical_cost": {"amount": 25, "currency": "EUR", "source": "search"}}
],
"legs": [ // filled by router, never by the LLM
{"from": "P123", "to": "P456", "mode": "foot",
"duration": "00:12", "distance_m": 900, "geometry_ref": "G1"}
]
}]
}, {
"id": "PL2", "name": "Portofino", "centroid": [...], "bbox": [...],
"vibe": "beach + seafood, one slow day",
"dates": ["2026-09-17", "2026-09-19"], "days": [...]
}],
"transit": [ // external transport: opaque, not OSM-routed
{"type": "train", // flight | train | bus | car_rental | hotel | ferry …
"from_place": "PL1", "to_place": "PL2",
"date": "2026-09-17", "depart": "08:10", "arrive": "09:55",
"ref": "Trenitalia 12345", "source": "user_booking",
"status": "booked"}, // booked|planned|estimated — booked = hard anchor
{"type": "car_rental",
"place": "PL2",
"pickup": {"at": "2026-09-18T15:00",
"where": "Hertz, via Palazzuolo 12", "poi_ref": "P901"},
"dropoff": {"at": "2026-09-21T10:00", "where": "…"},
"ref": "Hertz conf. ABC123", "source": "user_booking",
"status": "booked"}
],
"stays": [ // lodging = per-night records, NOT a trip attribute
{"id": "ST1", "place": "PL1",
"hotel": {"name": "Hotel Palagio", "poi_ref": "P123"},
"check_in": "2026-09-12", "check_out": "2026-09-15",
"ref": "conf. H-77821", "source": "user_booking",
"status": "booked|candidate", // candidates ⇒ worst-case anchoring for those nights
"price_per_night": {"amount": 180, "currency": "EUR", "source": "booking"}},
{"id": "ST2", "place": "PL2", "hotel": {"name": "hotel Y", "poi_ref": "P555"},
"check_in": "2026-09-17", "check_out": "2026-09-19",
"status": "booked"}
]
// router legs (intra-place) may themselves be multi-modal:
// {"mode": "multimodal", "sub_legs": [{"mode":"foot",…},{"mode":"bus",…}], …}
}
```
Rules that make this work:
- **Hierarchy is data, not ceremony.** A place is any `{label, centroid,
bbox, vibe, dates}`. "Italy", "Florence" and "Oltrarno" are the same
record shape; the planner just works at the granularity that has days.
One country trip = one place; a neighbourhood focus inside it = a child
place with its own days. No fixed category taxonomy anywhere.
- **External transit is opaque, and bookings are hard anchors.**
Flights/trains/long hauls are *not* OSM-routed (rail/air networks aren't
in OSM at usable fidelity): they're first-class `transit` objects with
`{type, window, source, status}`. Anything the user has *booked* (dropped
in as a ticket/confirmation — see Booking ingestion below) is a **hard
constraint**: the planner re-anchors days and legs around it and may
never move it, only work around it. `car_rental` objects additionally
*widen available modes* (car legs allowed inside the pickup→dropoff
window at that place). Intra-place movement stays in the router —
including **multi-modal legs** (walk + bus/subway; OSM carries the transit
network, GraphHopper/OTP/Valhalla compute the combined leg with sub-legs).
The validator checks *junctions* (arrive PL2 09:55 → first stop 10:30
feasible; transit departs after last stop; car legs only inside the
rental window).
- **Lodging is a set of *stays*, not a trip attribute.** A stay is
`{hotel, place, check_in, check_out, status, price_per_night}`. A day's
anchor *resolves* to the stay(s) covering its date — one city, one hotel
= one stay; a multi-city trip = several stays, each anchoring its own
dates. Multiple **candidate** stays covering the same dates are normal
(comparing two hotels before deciding), and they trigger the rule below.
- **Slots & states.** Every stop belongs to a **slot** (a free-form
mutual-exclusion group per day: `lunch`, `dinner`, `afternoon_visit`… —
not a fixed taxonomy) and a **state**: `planned` (counts against the
day's time budget), `maybe` (held, time-neutral), `alt` (candidate for
the same slot, rendered as an "or"). **One `planned` stop per (day,
slot)**; choosing a second demotes the incumbent to `alt` (swap-in
is one tap, nothing is lost). Removing the last `planned` of a slot
auto-promotes a remaining `alt`. A stop's **kind** is `point` (decided
POI), `region` (vague — "lunch near Oltrarno"; reserves the slot and
time, anchored to a centroid+radius), or `transit` (booked external
transport, a hard anchor). Region stops refine **in place** (the
concrete stop replaces the placeholder at the same position).
- **Worst-case anchoring (never plan on false precision).** Any leg whose
endpoint is a *set* — multiple candidate stays on that date, or a region
stop — is budgeted at the **farthest/worst member** and flagged in the
low-confidence tier ("worst-case until you commit/refine"). Committing a
hotel or refining a region recomputes those legs to real values; because
downstream decisions were made on worst cases, the plan stays valid.
Days end with a **return-to-stay** leg (included in the budget) except
on departure days.
- **Exclusion list.** "Not for me" / "no <name>" writes to the
`preferences.rules_out`-style list; it is applied to every future search
(query level, before anything renders) and fed to the LLM as context.
### Booking ingestion (drop a ticket, the plan adapts)
Users drag & drop real-world documents — train tickets, flight
confirmations, car-rental emails (PDF, image, screenshot) — onto the app:
1. **Parse (LLM):** document → structured proposal
`{type, date, times, origin, destination, reference, carrier}`. The LLM
extracts; nothing is invented — fields the document doesn't state stay
null and are asked for in the UI.
2. **Ground:** origin/destination names are geocoded to POIs/places; the
proposal is matched against the itinerary (which place, which day, near
which stop).
3. **Propose:** rendered as a *suggested transit object* (status
`booked`) plus the planner's re-anchoring diff: which legs/days move
around it, new mode availability (car), junction checks. Example: drop
a rental-car confirmation "pick up 15:00, via Palazzuolo" → car appears
available from 15:00, afternoon legs re-plan onto car, walking leg to
the agency inserted.
4. **Confirm:** user accepts (commits via `apply_edit`) or corrects fields
inline; the document's reference number is stored on the object and
shown on the itinerary.
This is the mechanism by which the itinerary *subsumes the user's real
commitments* instead of competing with them — the app plans around what's
already booked, which is the state of the world on day 1 of any trip.
**Hotels are special bookings:** a lodging booking sets each day's `base`
— the anchor that mornings depart from and evenings return to (drives
first/last leg computation and "open near your hotel" suggestions). Hotels
also enter the alternatives loop ("a quieter hotel near the old town?"),
and changing the hotel re-anchors every affected day.
**Entry via booking:** the app's first screen is the "Where would you like
to go?" prompt — and it doubles as a drop zone. Dragging in a booking
(train/flight most common, hotel reservation works too) starts the app:
parse → ground → the itinerary is born already anchored to the user's real
commitments, with the places implied by the booking pre-selected.
- **The LLM proposes, the tools decide.** The LLM adds/moves/removes *stops*
and *places* (POI refs, vibes, date ranges). Every *leg* (mode, duration,
distance, geometry) is computed by the router and written back by a
service, never by the LLM. Arrival/departure times are *proposed* by the
LLM and *validated* by the checker (below); the checker can reject or
correct ("Museum A opens 10:00, your 09:30 arrival is impossible — moved
to 10:00").
- **Deterministic validator after every edit** (`validate_itinerary`):
- opening-hours check of every stop against its scheduled time,
- time-budget feasibility (arrive ≤ depart, legs sum to the day),
- mode/time sanity (no 2 a.m. entries in closed venues, max walking
distance per leg, etc.).
The validator's warnings are fed back to the LLM and/or the user. This is
the itinerary equivalent of the grounding rule for plain QA.
- **Versioned diffs → the refinement UX.** "Make day 2 more leisurely",
"swap the gallery for a pub with live music", "we're staying 1 day longer"
become edits that produce a *new version*; the UI renders old vs new
(map + timeline) with accept/reject/undo. Refinement is where the product
lives — cheap incremental edits, not full replans.
- **Persistence & shareability:** itinerary is a JSON doc in Postgres (+ map
payload); exportable/shareable; re-validation is a pure function so a
shared itinerary can be re-checked when data refreshes.
### Suggestion engine: knowledge sources
Suggestions for places/vibes/stops are a **three-source blend**, in order of
latency and trust:
1. **OSM ground truth (instant, local):** `place_stats` — what the place
actually contains (tag clusters, coastline, parks, transit). Sets the
floor: nothing is suggested the place can't deliver.
2. **LLM world knowledge (instant):** culture, seasonality, "what's worth
it and why", rough quality rankings, hidden gems. Fast but may be
stale or wrong.
3. **Web search (seconds, server-side):** augments 1+2 — current events
("restored 2026", "festival in September", closures, new openings),
practical detail ("is the cable car running"), **per-stop expectations**
(recommended dwell time, entry price, typical meal duration & cost —
see Stop expectations below), and **freshness checks** on anything the
LLM asserts confidently. Implementation: a search backend behind one
interface (self-hosted **SearXNG** for privacy/free, or Tavily/Serper-
class APIs for quality); queries derived by the LLM per suggestion,
results cached per (place, topic, week), and **cited** — each
search-derived enrichment carries its source URL, shown in the UI.
The blend is per-suggestion, not global: a "visit X" card may be 100% OSM+
LLM (no search needed), while "best time to see the aurora here" is
search-heavy. Search **augments; it does not gate** (below).
### Real-time, non-blocking UI (streaming enrichment)
Search is seconds-slow; the chat and cards must feel instant. Pattern:
- **Two-phase rendering per turn.** The LLM's fast-path output (OSM +
world knowledge) renders immediately: vibe chips, suggestion cards with
name/pitch/vibe, draft edits. Each card is *fully interactive the moment
it appears* (tap to like/rule-out/add to plan).
- **Async enrichment over SSE.** While the user reads, search results
stream back (server-sent events, one channel per conversation turn) and
*patch* existing cards in place: added image, "opens Sep 2026", source
link, a corrected fact. Enrichment never reorders, removes, or rewrites
committed content; if a search result contradicts a rendered claim, the
card shows the correction inline (note + source) and the next LLM turn
sees the corrected state. User actions made while enrichment is in
flight are valid — the server resolves conflicts against the itinerary
version, not the search results.
- **Progressive depth in chat:** the assistant answers in layers — direct
answer first, then "digging deeper…" details as search lands. The user
can steer or cancel at any time; a fast reply the user acted on is never
retracted silently.
### Imagery in the UI
Pictures guide activity choices (a beach, a monument, a valley). Sourcing,
cheapest-first:
1. **Wikimedia Commons** (free API, no key, per-image license +
attribution) — keyed off the place/POI name; covers monuments, nature,
cities well.
2. **POI-level photos from OSM tags** where present (`image`/`photo` tags).
3. Optional paid upgrade: Unsplash/Pexels API for consistent lifestyle
shots.
Rules: an image is a *card decoration with attribution + source link*,
fetched server-side and cached (keyed by POI/place), lazy-loaded, with a
graceful fallback (category icon over a gradient) when nothing is found —
an empty image slot must never block card interaction.
### Exploration & preference loop (the fluid "where would you like to go?")
Idea generation is a *conversation*, not a form. The `preferences` block
above is the state of that conversation, and it evolves through two input
channels: chat text and **UI selections** (liking/ruling out suggestions,
picking a vibe, dragging a suggested stop into a place).
- **Seeding ("what is this place like?") is data-grounded.** When the user
names a destination, the server pulls its OSM *tag statistics* (top
tourism/amenity/leisure clusters: beaches, trails, museums, nightlife,
hiking, food culture; plus coastline/park area) — a small deterministic
function — and the LLM turns that into **vibe suggestions** ("relaxed",
"beach", "adventure", "food crawl") *for that specific place*, with web
search (Suggestion engine above) asynchronously adding current events
and seasonality ("harvest festival in September") as chips enrich. This
kills the failure mode of suggesting "beach days" for a landlocked
region: the chips are generated from what the place actually contains,
with the LLM adding world knowledge on top.
- **Suggestion loop:** the LLM surfaces *candidate* places/days/stops in the
UI as cards (not just chat text). Every user reaction — like, pass,
"not that, something quieter", a free-text refinement — writes into
`preferences` and the LLM regenerates/re-ranks. This is the same mechanism
as the alternatives loop, applied upstream: *generation* is iterative
constrained sampling, and the constraint set is user-observed, never
invented.
- **Progressive commitment:** suggestions are `candidate` → user picks →
`tentative` → planner fills in → `confirmed`. The trip takes shape top
down (country → places → days → stops) but any level can be edited at any
time; committing deeper doesn't lock higher levels.
- Free-form refinement field: always on. Its text goes to the LLM with the
current `preferences` + current suggestion batch; the LLM's reply may
include both prose and new suggestion cards. Nothing here requires
predetermined categories — the only structured bits are the preference
keys, which are open strings.
### Stop expectations: dwell times, prices, meals
A bottom-level stop is more than a slot: it carries **expectations** that
the plan is checked against:
- `planned_duration` (what's scheduled) vs `suggested_duration` (how long
it *should* take, with `source: user|osm|llm|search`).
- `price` `{amount, currency, basis (per person/family), source}` — OSM
`fee`/`average_spend` tags first, else background search, else labeled
LLM estimate.
- Meal stops add `meal {kind, typical_duration, typical_cost}`.
- **User overrides are sticky and authoritative.** If the user sets 1 h at
the Palazzo Pitti, `source: user` — the LLM never silently re-extends it;
at most a one-time inline note ("guides usually suggest 23 h"), never
nagging. The UI always shows planned vs suggested ("1 h (suggested ~2.5
h)") so the trade-off stays visible without being pushy.
- **Background search fills missing expectations** through the enrichment
pipeline (Suggestion engine §3) whenever a stop is added — dwell, price,
meal duration & cost — streaming into the card like any other
enrichment, cited.
- **Meals are a first-class planning target, not an afterthought:**
- `purpose: meal` stops carry a kind (breakfast/lunch/dinner). After an
activity is confirmed, `suggest_meals(day, window, near)` actively
finds restaurants: within a short walk of / along the corridor between
the neighboring activities, open during the free window, whose
`typical_duration` + travel fits the gap — each with cost and a
one-line character note.
- Day budgeting books meals at *typical* durations (lunch ≈ 6090 min,
not 45), so free time isn't silently eaten by eating.
- Prices are collected per stop and rolled up in the background (per
day and trip, by category) — **not surfaced in the UI at this
stage**: no cost totals, no budget bars, no cost review issues. The
data model keeps it ready to surface later without rework.
### Plan review (overall & detailed — issues, not prose)
Layered, mirroring hard/soft:
1. **Deterministic validator (existing):** opening hours, feasibility,
booked anchors, mode windows — blocking.
2. **Expectation-aware checks (new, deterministic):** planned vs suggested
dwell (45 min booked for a meal typical at 90 → issue), price vs budget,
daily walking distance, dead gaps, last-stop vs closing times.
3. **LLM review — `review_plan(scope: trip|place|day|stop)`:** reads
`itinerary_summary` + the computed expectation deltas and returns
**structured issues, not prose**: `{severity: error|warn|info, scope,
stop?, kind (meal_too_short | pacing | fatigue | cost | hours_risk | …),
message, fix_op?}`. The LLM catches what heuristics can't (three
museums in a row, dinner at 21:30 with the last train gone, a brutal
transition day after a lazy one); the narrative lives in `message`.
Issues are a **UI object**: an issues panel (filter by scope/severity;
click → zoom to the stop; one-click apply of the attached `fix_op` via
`apply_edit`). Review runs on demand ("review day 2") and automatically
after large edits (badge lights up when new warnings appear — never
blocks). Every expectation cited in an issue carries its source, so the
user can see whether "typically 90 min" came from search or a guess.
### Itinerary tools (added to the catalog below)
- `place_stats(location | place_ref) -> tag statistics`
Deterministic: OSM amenity/tourism/leisure clusters, coastline/park area
for the place's bbox. Input to vibe generation — never invented by the LLM.
- `suggest_vibes(place_ref)` / `suggest_places(dates, budget)` — LLM-facing
wrappers that emit **structured suggestion cards** (place: name, label,
1-line pitch, vibe, suggested days; stop: POI ref, pitch, trade-off).
Cards render in the UI; user reactions are recorded *server-side* into
`preferences` (the LLM never has to "remember" preferences).
- `add_place(name | bbox, dates, vibe)`, `remove_place(ref)`,
`move_place_dates(ref, dates)` — hierarchical edits; adding a place
auto-creates a `transit` stub from the previous place.
- `suggest_meals(day, window, near_stop?) -> [restaurant cards]` — nearby
restaurants fitting the free window (open at that time, typical_duration
+ travel fits), each with cost, typical duration, hours, one-line
character note; expectations background-filled by search.
- `review_plan(scope: trip|place|day|stop) -> {issues[]}` — deterministic
expectation-aware checks + LLM review; issues carry severity, kind,
message, and an optional one-click `fix_op`. See Plan review.
- `transit_estimate(from, to, date, mode?) -> {windows, duration, source,
cost?}` — external rail/air lookup where available, else labeled estimate;
never fabricated silently — `source` is always shown in the UI.
- `plan_day(constraints, candidate_pool?) -> draft day` — server-side
heuristic planner (geo-cluster + hours + mode) that proposes a draft; the
LLM reviews/edits instead of freehand-sequencing from scratch. Keeps the
LLM's job in its comfort zone (taste/judgment) and the combinatorics in
code.
- `add_stop(day, poi_ref, slot?)`, `move_stop(...)`, `remove_stop(...)`,
`replace_stop(old_ref, constraints)` — atomic edits; each triggers leg
recomputation + validation.
- `validate_itinerary(itinerary) -> {ok, warnings[], fixes[]}`
- `itinerary_summary(itinerary) -> compact text` — the context form the LLM
sees (POI names, slots, leg durations; not geometries), keeps the context
window small on long trips.
- `parse_booking(doc) -> transit proposal` — LLM extraction from a dropped
ticket/confirmation (PDF/image) → structured fields (nulls asked for, not
invented) + geocoded origin/destination; feeds the Booking ingestion
flow.
- `apply_edit(op)` — **the single mutation endpoint for ALL itinerary
changes, from chat or from the UI** (both emit the same op vocabulary:
`move_stop`, `set_slot`, `set_leg_mode`, `set_route_via`, `add_stop`,
`remove_stop`, `add_transit`, …). Applies op → recompute affected legs →
validate (booked transit = immovable anchors) → version bump → stream
new state. Shared undo/redo lives here.
- `recompute_legs(day)` — router recomputation after edits (also called
implicitly by `apply_edit`).
- `web_search(topic, place_ref?, freshness?) -> [results with source URL,
date]` — the search backend (self-hosted SearXNG, or Tavily/Serper-class
API behind the same interface); per-turn budget (e.g. ≤4 queries),
cached per (topic, place, week).
- `image_for(poi_ref | place_ref | query) -> {url, attribution, license}`
— Wikimedia Commons → OSM photo tags → fallback icon; cached.
The existing QA tools (geocode/search/hours/route) are the *building-block*
tools the itinerary tools are composed from, and they also serve ad-hoc
questions ("anything good near the castle?").
## Planning lifecycle (functional scope)
Five capabilities, each mapped to the machinery above:
| Phase | User does | System does | Key tools/mechanisms |
|---|---|---|---|
| 0. **Exploration** | "where would you like to go?" + free-form refinement; like/rule-out suggestions | OSM-grounded vibe suggestions per destination; candidate *place* cards; every reaction updates `preferences` | OSM tag-statistics function + LLM, suggestion cards, preference state (see Exploration loop) |
| 1. **Idea generation & filtering** | "we like history and coffee" | proposes candidate POIs (broad recall), surfaces hours/cuisine/tags, filters by constraints & interest match | `search_poi`, semantic vector path, `poi_details`, candidate pool with `status=candidate` |
| 2. **Organization & planning** | "plan day 2 around these" | clusters geographically, sequences stops, **multi-modal route planning** (foot/bike/car/train legs, with transfer stops), schedules slots | `plan_day`, `route` (per-mode legs), `distance_matrix`, validator |
| 3. **Changes & refinement** | "move the museum to the morning" | minimal-diff edits, re-derives affected legs & slots, re-validates | `move_stop`/`add_stop`/`remove_stop`, `recompute_legs`, versions |
| 4. **Alternatives** | "anything better for lunch?", "what if it rains?" | generates ranked *alternative sets* (25 options with trade-offs: distance, hours, price, weather-fit) for the user to pick from — the LLM presents, the user decides, the choice commits one option | `search_poi` + `distance_matrix` + `opening_hours` composed into an `alternatives` payload rendered as selectable cards |
| 5. **On-trip execution** | "where next?", "open right now?" | serves the synced itinerary, live re-validation (hours, traffic mode), small edits | read of synced itinerary, `opening_hours`, re-route |
Notes:
- **Multi-modal legs:** a leg is `{from, to, mode}`; the router returns the
best single-mode leg, and a separate *transfer planner* composes
multi-leg chains (e.g. train + walk) with intermediate stops. Transfer
chains are stored explicitly so refinement can swap a leg's mode without
touching its endpoints.
- **Alternatives are first-class UI, not chat text:** an alternatives request
produces a structured set (options + one-line trade-off rationale each);
the user taps one and it commits. This keeps selection a *human* decision
and keeps the LLM out of the final pick.
- **Scope honesty for multi-country:** intra-place = OSM routing + hours +
POIs (strong); inter-place = opaque transit objects (estimates or external
rail/air APIs) — the system plans *junctions and day structure* across
countries, it does not pretend to book or exact-route the long hauls.
## Data layer
Source of truth: **OpenStreetMap** (free, structured, has `opening_hours`,
`brand`, `amenity`, `shop`, `cuisine`, `opening_hours` etc. per POI).
| Component | Options | Notes |
|---|---|---|
| Raw map data | Geofabrik PBF extract for the region of interest | one-time download + periodic refresh |
| Spatial DB | **PostGIS** (Postgres) | import PBF via `osmium`/`osm2pgsql`; gives SQL spatial queries (radius, nearest, bbox) |
| Geocoding | **Nominatim** (local) or Photon | address ↔ coordinates; fuzzy place names |
| Routing | **GraphHopper** (easy multi-modal) or **OSRM** (fast, table API) or **Valhalla** | car/bike/foot/wheelchair profiles, distance + duration, turn instructions |
| Opening hours | `osm-opening-hours` lib (Go/Py/JS) | evaluates OSM `opening_hours` for "now" / "at time X", timezone-aware |
| (optional) live queries | Overpass API | for regions we don't keep locally; rate-limited |
Why local over pure APIs:
- reproducible, fast, no rate limits, works offline;
- PostGIS lets the LLM's tools do real spatial math (e.g. "shops within
500 m of the route"), which no hosted API offers as a single call.
Alternative for v0: skip PostGIS entirely and use Overpass API + a hosted
router (GraphHopper API / OSRM demo). Fine to start, but keep the interface
so a local backend can be swapped in.
**Multi-region note:** with multi-country itineraries the data unit is
*per place* (per-region PBF / router graph / tile cache), not one global
load. Local per-city GraphHopper is fine; alternatively a hosted router
serves any place. Inter-place transit is **not** OSM routing — it's external
(rail/air search APIs where they exist, e.g. open rail timetables, else
labeled LLM/user estimates) stored as opaque `transit` objects.
## Architecture
```
┌────────────────────────────────────────────────────────────┐
│ Frontend: chat UI + map view (Leaflet / MapLibre GL) │
│ (renders POIs, routes the LLM refers to; click → context) │
└──────────────┬─────────────────────────────────────────────┘
┌──────────────▼─────────────────────────────────────────────┐
│ Agent layer (the LLM, chat + function calling) │
│ - conversation state (user location, preferences) │
│ - itinerary state (current version, constraints) │
│ - tool selection & multi-step planning │
│ - answer synthesis with citations (POI ids / coords) │
└──────────────┬─────────────────────────────────────────────┘
│ tool calls (JSON)
┌──────────────▼─────────────────────────────────────────────┐
│ Tool / service layer (deterministic, no LLM) │
│ geocode · search_poi · poi_details · nearest · route · │
│ distance_matrix · opening_hours · along_route · │
│ route_options · corridor · stop_cost · optimize_stops · │
│ place_facts │
│ ┌────────────┬──────────────┬──────────────┬───────────┐ │
│ │ PostGIS │ Router │ Geocoder │ Hours lib │ │
│ └────────────┴──────────────┴──────────────┴───────────┘ │
└────────────────────────────────────────────────────────────┘
```
### Tool catalog (the heart of the system)
Each tool = a JSON-schema function the LLM can call. Keep them narrow and
composable.
1. `geocode(query) -> [name, lat, lon, type]`
Resolve "St. James Park", "my location", street numbers.
2. `search_poi(name?, category?, tags?, near: {lat,lon,r} | bbox, limit)
-> [pois]`
Category synonyms handled in the LLM prompt ("place to get coffee" →
`amenity=cafe`). Returns id, name, tags of interest.
3. `poi_details(poi_id) -> full tags`
Second-level lookup so list queries stay cheap.
4. `nearest(category, from, max_results, max_distance)`
5. `route(from, to, profile=car|foot|bike|multimodal, via?, depart_at?) ->
{distance, duration, geometry, steps, sub_legs?}` — `multimodal`
(walk+bus/subway, from OSM transit data) returns sub-legs; `via[]`
waypoints make user-modified routes first-class (drag a waypoint,
recompute, done). The same call backs the L3 route editor with live
duration updates
6. `distance_matrix(places[], mode) -> matrix`
For "which of these is closest / quickest".
7. `opening_hours(poi_id | poi_ref, when=now|ISO-timestamp) ->
{open?, closes, weekly}`
8. `along_route(route_ref, category, max_distance) -> pois`
"a petrol station along the way" — PostGIS buffer over route geometry.
*Driving trips upgrade this to corridor search — see below.*
9. `reverse_geocode(lat, lon)` — for "what is near here".
10. `route_options(from, to, profile, max_alts) ->
[{ref, geometry, duration, tolls, via: [highways], diff_vs_fastest}]`
Fastest + user-relevant alternatives. A driving trip **must not emit
any time estimate before this has run** — the route is computed, never
narrated from LLM memory ("you'll cross Boston at 15:15" is a bug,
not an answer).
11. `corridor_from_route(route_ref, buffer_km, time_window?) -> corridor`
Spatial predicate (band around the geometry) **+ an along-route
coordinate system** (fraction / minutes-from-origin). User constraints
— "south of origin", "in the first 2 hours", "on the way" — are
filters in this coordinate space; direction is a fact, not a guess.
12. `search_along_route(corridor, slot, tags, prefs) ->
[{poi, along_route_pos, detour_min, facts}]`
OSM POIs inside the band LLM world-knowledge candidates *projected
into the band* (a remembered trail outside the corridor is filtered or
explicitly flagged "off route, +X min"). Exclusion list applied.
Every candidate carries its computed detour cost — the sort key.
13. `stop_cost(route_ref, stop | poi) -> {detour_min, total_added_min,
new_route_ref}`
`(route A→B via stop) (route A→B) + dwell + declared overhead`.
The single most important number in trip planning: pure router
arithmetic, no LLM. Surfaced as a chip ("+18 min to trip") and used
by the validator (arrival-time constraints).
14. `optimize_stops(route_ref, candidates[], k, time_budget) -> [ordered
stops]`
Minimizes detour + Σ dwell + Σ overhead over ordered subsets of
exactly k (exhaustive permutation search ≤ 10 candidates — also the
benchmark ground truth — else greedy best-insertion). "Pick my two
45-min hikes" is *this* call, not an LLM freehand choice — the LLM
frames candidates/tastes, the optimizer picks.
15. `place_facts(poi_id) -> {distance, loop_length, elevation_gain,
est_duration, dog_policy, parking, season_notes, hours, provenance[]}`
Sourced activity facts (OSM tags + cached source pages). The answer
may only quote numbers that exist here.
Itinerary tools (`plan_day`, `add/move/remove/replace_stop`,
`validate_itinerary`, `itinerary_summary`, `recompute_legs`) — see the
Itinerary section above.
Design rules for tools:
- Return compact JSON; truncate large fields; include a stable `id`/
`map_ref` per POI so the frontend can highlight it and the LLM can refer
to it without repeating all the data.
- Every tool returns coordinates so the answer is always renderable.
- Time is explicit: `now` resolved server-side with the location's
timezone (IANA), passed through opening-hours evaluation.
- **Provenance is enforced, not stylistic.** Every time/cost/policy value
carries a source tag: `computed` (router) / `measured` (OSM) /
`sourced` (booking, official page) / `assumed` (LLM). A hard output
rule: *an assumed number may never be presented as fact* — it either
triggers a tool call or renders in the low-confidence tier. (Failure
case that motivated this: a real planning session in which the route,
distances and a 40-mile "detour" were all confidently asserted from LLM
memory, and the "detour" turned out to be on the actual fastest route.)
### Driving trips: the route is a first-class object
Requirements sharpened by a real exercise (Lincoln NH → Queens NY,
"two 45-min hikes on the way") where the LLM-only approach failed on
route, direction, distance and activity facts while the *activity
recommendations themselves* (given correct spatial facts) were good. The
lesson: **bad facts poisoned good judgments**, and the fix is structural:
1. **R1 — Compute the route before speaking.** A driving trip has an OD
pair; `route_options` runs first, the user's route choice (fastest vs
scenic vs toll-avoiding) becomes the versioned `route` object, and the
corridor derives from it. No suggestion, ETA or "you'll cross X at Y"
may be emitted before a route exists.
2. **R2 — "Along the route" is a spatial primitive.** Search is
corridor-constrained by default; *direction* and *position* constraints
("south of origin", "first 2 hours") are filters on the corridor's
along-route coordinate system, not LLM geometry.
3. **R3 — Every stop has a computed cost.** `stop_cost` (detour + dwell +
overhead) is the display number, the sort key, and the input to the
validator ("detour pushes past your 19:00 arrival"). Adding/removing
stops re-derives the route and all costs through the existing mutation
pipeline.
4. **R4 — Provenance-gated output.** See design rules above.
**Model impact:** the itinerary gains a versioned `route` per place/trip
(geometry ref, profile, chosen-alternative); stops gain `corridor_pos`
(minutes/fraction along route). The validator checks stop costs against
user time constraints ("home by 19:00"). **UI impact:** stop cards get a
`+X min` detour chip (computed tier); focus-mode pins are corridor-
constrained by default; the budget bar counts detour minutes, not just
dwell.
**Evaluation hook:** *corridor-aware stop optimization* is a cleanly
bench-markable task class: given OD + k + time budget, ground truth =
router-optimal stop set and total time; compare system output. Generate
the benchmark from real OSM + router data (TravelPlanner covers leg
planning but not this composite). **Built:** `router/` (Go, stdlib-only) —
`route_options`/`stop_cost`/`optimize_stops`/corridor over Valhalla or
local OSRM, `cmd/bench` with 5 real NE-corridor tasks and recorded
goldens (26 checks passing), integration tests asserting e.g. the
fastest Lincoln NH → Queens NY drive stays ≥15 km from Boston.
**`optimize_stops` semantics (implemented):** the user asked for *k*
stops, so the objective is total = detour + dwell + overhead over
ordered subsets of **exactly k**. "At most k" without a budget
degenerates (zero stops always wins — dwell only adds cost). With a
time budget: best feasible exactly-k → best feasible smaller set
(flagged `relaxed`) → best exactly-k ignoring the budget (flagged
infeasible). Exhaustive permutation search ≤ 10 candidates (this *is*
the ground truth), greedy best-insertion above.
### Retrieval strategy (hybrid)
- **Structured path (primary):** tools 19. Exact, cheap, verifiable.
- **Vector path (secondary):** embed the POI text (name + tags + any free
text) once per POI, index in Postgres `pgvector` (or SQLite-vec for v0).
Used for fuzzy semantic queries the tag vocabulary can't express:
"a place that sells artisan sourdough", "quiet pub with a garden".
Semantic search returns candidates → `poi_details` → LLM decides.
- Rerank candidates by distance (PostGIS) before handing to the LLM so the
context window always holds the *nearest, most relevant* few, not ten
random hits.
### Agent loop
Standard function-calling loop, with the itinerary state injected as
context (via `itinerary_summary`, not the raw doc):
```
user question
→ LLM (system prompt: you are a map assistant; user location; tool list;
answer only from tool results; cite POIs; give times as durations,
don't invent)
→ tool call(s), possibly parallel (geocode A, geocode B)
→ results back → more calls or final answer
→ answer + structured payload {pois:[refs], route: {geometry}}
```
- **Grounding rule:** the LLM must never state a travel time/distance/hours
value that didn't come from a tool result. This is the main guard against
hallucination and the reason the architecture matters.
- Multi-hop questions work naturally: "Is the best-rated bakery near my
home open right now?" → `search_poi``route` (optional) →
`opening_hours` → answer.
- Cap the loop (e.g. 6 tool rounds) and fall back gracefully.
- **Two operating modes:** *QA* (question → answer, no state change) and
*refine* (answer **plus** an itinerary edit). In refine mode the turn
commits as a new itinerary version; the validator runs between the edit
and the user sees anything.
### Interface (web-first; mobile on-trip)
Quality bar: **beautiful, responsive, flexible** — the map is the hero, the
chat is a tool, and every LLM suggestion must be visually inspectable and
reversible.
**Web app (primary, desktop-optimized):**
- Layout: **map (hero, ~6070%) + day timeline rail + chat pane** (collapsible
to a bottom drawer on narrow screens). Responsive down to tablet/phone —
same app, panels reflow; it is *not* the on-trip app.
- Timeline rail: days → stops in sequence with slots, leg chips (mode
icon + duration); drag a stop to reorder (triggers `recompute_legs` +
validator), hover a leg to highlight its geometry.
- Alternatives & candidates render as **cards** (photo/tagline, hours, trade
-offs) alongside the map, with one-tap "put in plan".
- Every agent turn carries a **structured payload** (POI refs, geometries,
itinerary diff). Refine-mode edits render as an **accept/reject/undo diff**
(map + timeline animate old→new); rejection feeds back to the LLM as
conversation context.
- Power features for desktop: compare two versions, pin a POI as a
conversation anchor, filter layers (hours, price, mode), export
(PDF/ICS/GPX), split-screen two days.
- **Issues panel:** `review_plan` output lives here — filter by
scope/severity, click an issue to zoom to its stop, one-click apply of
the suggested fix-op; badge lights up when an edit introduces new
warnings (never blocks). Stops show planned vs suggested time ("1 h
(sug. ~2.5 h)") and price chips with source attribution.
**Focus mode (slot search — the Google-Maps pattern).** "Where should we
lunch? What for the afternoon? Which hotel?" is a *different mental mode*
from editing the plan, and the UI says so:
- Triggered by chat — "find a lunch spot", "lunch near Oltrarno",
"alternative hotels", "find another activity" — and works whether the
slot is **empty, vague, or already decided**: search is always
**re-openable**, and choosing is a *swap* (incumbent → `alt`, chosen →
`planned`), so any decision can be revisited later.
- The rail dims; the map becomes the hero with **name-labeled teardrop
pins** on the candidates and nothing else. Click a pin → popup: pitch,
**distance from the day's anchor (the previous stop — where the user
will actually be) and to the hotel**, duration, price, actions.
- **Chat becomes the filter bar**: "cheaper", "near Boboli", "views",
"vegan", "no Mario" — pins re-filter live; each reply names the closest
remaining option to the anchor. Filters are slot-aware ("cheaper" means
≤ €25 for lunch, ≤ €220 for a hotel).
- Actions per pin: **Choose this** (slot swap), **Maybe**, for hotels
**Keep as option** (adds a candidate stay for those nights ⇒ worst-case
anchoring) and **Make this the hotel** (commit ⇒ re-anchor, real legs),
and **More like this** (re-query on the candidate's own tags).
- Candidates come from the 3-source engine (OSM ground truth → LLM world
knowledge → web enrichment); the **exclusion list** is applied on entry.
**Mobile (web-first, usable from day one).** Below ~860px the 3-pane grid
collapses to a **single pane with a bottom tab bar (Plan | Map | Chat)**:
each pane full-width, the topbar wraps to two rows with horizontally
scrollable day tabs, and the map re-fits to the full viewport. Same app,
same state — not a separate mobile build. (A native on-trip shell with an
offline trip bundle is a later phase, not a prerequisite.)
**Direct editing is first-class — chat is one editor, not the only one.**
Anything the LLM can do to the itinerary, the user can do by hand, and both
paths funnel through the **same mutation pipeline**:
- Every change — LLM or human — is an atomic *edit op* (`move_stop`,
`set_slot`, `set_leg_mode`, `set_route_via`, `add_stop`, …) sent to the
server's single `apply_edit` endpoint. The server recomputes affected
legs, runs the validator, bumps the version, and streams back the new
state. The LLM never gets a fast path around validation, and a user drag
never bypasses leg recomputation.
- Shared **undo/redo** stack across both editors (each op is a version
delta); an LLM suggestion the user rejects is just an unapplied op.
- Hand-editing affordances: drag stop reorder / drag across days; inline
slot-time editing (typing a time or dragging on a time ruler — later
slots cascade and the validator flags overruns live); per-leg mode
switcher (foot/bike/car) with instant duration update; **route editing:
drag waypoints on the map (or click to add them) to alter a leg's path —
for a hike, snap to the actual trail — and duration/distance recompute
live**; add a stop by clicking the map (reverse-geocode → search → pick);
remove/duplicate stops; per-place day-length and budget settings.
- While the user is editing, chat suggestions for the same stops are
dimmed, not force-updated (the UI never fights the user's hand).
**Time is the product's core value — a first-class display object, not
metadata.**
- Every stop shows its *duration*; every leg shows *mode + duration*
(distance on hover); each day shows *planned total vs available* as a
budget bar (7.5 h planned of 11 h waking; amber/red on overrun); places
and the whole trip roll the totals up.
- **Travel times are never LLM-estimated.** Leg durations come from the
router (mode-specific profiles) or from explicit `transit` objects
(scheduled or labeled estimate). *Dwell* times (how long at the museum)
are a different beast: they carry `planned` vs `suggested` with source
labels, and user-set values are authoritative (see Stop expectations).
All estimates render with a "~" (dashed chip), distinct from computed
values.
- **Confidence hierarchy — every time/number estimate carries a level, and
the UI encodes each level distinctly:**
| Level | Source | UI indicator |
|---|---|---|
| 5 (top) | user-set | filled chip + pin/lock glyph |
| 4 | scheduled (booked transit / hotel) | filled chip + ticket glyph |
| 3 | router-computed | plain filled chip |
| 2 | search-derived | "~" + source link |
| 1 (lowest) | LLM estimate | dashed chip, labeled "estimate" |
Rules: a lower-confidence value never silently replaces a higher one
(user overrides are sticky — §Stop expectations); background enrichment
that *raises* confidence (1→2, 2→3) streams into the UI as an update
with a small "updated" pulse, never a silent rewrite.
- Time changes animate: a recompute that moves the afternoon is visible
(slots slide on the rail, the budget bar moves), so the user always sees
the *consequence* of an edit in minutes, not just in words.
**Multi-scale navigation (one UI, four zoom levels):**
| Level | View | Map shows | Rail shows |
|---|---|---|---|
| **L0 Trip** | all places + transit | places as nodes, transit arcs, day counts | places → days as a strip |
| **L1 Place / Day** (primary working view) | one place or one day | stops + intra-place legs | stops with slots + leg chips |
| **L2 Time window** | a slice of a day ("after lunch") | that subset | the subset, rest dimmed |
| **L3 Leg / Activity** | a single stop **or a single leg** | **the route editor**: the path with draggable waypoints (hike: snap to the trail; urban: walk-only, or multi-modal walk+bus+subway with the router's sub-legs shown as segments), distance + **live-updating duration** as the user edits it, optional elevation profile (GraphHopper elevation API) | day context, collapsed |
Drill down by click (place → day → time window → leg/stop), back via
breadcrumb; map zoom and rail scope stay in sync. L3 is the **generic route
editor** — a hike is just a foot-only instance of it; so is a walk through
the city, or a "walk + bus + metro" chain (the router returns the
multi-modal leg with sub-legs; the user can drag waypoints on any of them
or swap a sub-leg's mode and the duration updates live). It must feel
instant: leg recompute for short legs is tens of ms on the local router.
The mobile on-trip app uses the same hierarchy (L0L3) with direct editing
scoped to today (reorder, skip, route-edit) and read-mostly elsewhere.
**Mobile app (on-trip, secondary):**
- Modes: *today* (current day as a checklist timeline + map, live "open now"
badges), *navigate* (leg-by-leg with mode-aware guidance), *ask* (chat —
server LLM when online), *edit small* (reorder today's stops, drop a
stop; server re-plans legs when connectivity allows).
- **Offline trip bundle** — the key on-trip requirement (limited coverage
on the trip itself):
- On "prepare for trip" the phone downloads: the itinerary + full stop
POI details (hours, tags, contact), **vector map tiles** for the trip
bbox (MapLibre + locally generated tiles, e.g. via `osm2vsmartin`, or
raster MBTiles), and leg geometries with turn-by-turn instructions.
- Offline the app serves: itinerary, maps, stop details, hours evaluation
(the hours library runs fully offline — pure function of local time),
and walking/bike re-routing within the trip area if we ship a small
local router (GraphHopper can preload a small region; treat as a v2+
stretch — offline re-route degrades gracefully to "show saved leg").
- **Sync model:** bundle = snapshot (itinerary version N + tile version);
edits made offline queue and reconcile on reconnect (itinerary is a
versioned doc, so a simple last-writer + server re-validation works);
re-validation on sync flags stops whose hours/legs changed.
- No LLM offline (except the optional on-device 4B for "is X open?" —
see Model selection); the offline experience is *plan + navigate*,
not *re-plan*.
## Hard sub-problems to plan for
1. **Opening hours**: OSM coverage is patchy (especially for small towns);
answer honestly ("no hours listed") instead of guessing.
2. **Timezones**: evaluate hours in the *place's* timezone, show the user
their timezone.
3. **"Open now" at the route level**: `along_route` + `opening_hours`
composition is a differentiator.
4. **Ambiguity**: geocoding "Main Street" → ask a clarifying question or
show top-3 candidates in the UI (structured, not just text).
5. **Staleness**: schedule PBF refresh (e.g. weekly); show data timestamp.
6. **Offline bundles**: tile size vs bbox (vector tiles keep this small;
a city at z14 ≈ a few hundred MB raster, much less vector); offline-edit
conflict resolution (versioned doc + server re-validation is the simple
correct answer); bundle validity (re-validate on reconnect).
7. **Time is the value proposition**: calibrate router profiles per mode
where possible; always distinguish *scheduled* (transit), *computed*
(router) and *estimated* (LLM/user) times in the UI; L3 route-edit
recompute latency must stay <~100 ms for short legs (local router,
preloaded graph).
8. **Two editors, one truth**: chat and hand-editing share `apply_edit`
conflict resolution is by itinerary version (last op wins, validator
re-checks), never by "who asked first".
9. **Booking-document parsing**: tickets/confirmations arrive as PDFs,
screenshots, and carrier-specific layouts — LLM extraction needs
per-field confidence (show "did we read the platform right?" for
doubtful fields), and station/agency names need robust geocoding
("near via Palazzuolo" → closest matching POI with user confirmation).
Car-rental windows are easy to get subtly wrong (timezone, date line) —
always echo the parsed window back for confirmation.
## Evaluation
- **MapQA** (arXiv 2503.07871) — 3,154 QA pairs from OSM (Southern
California, Illinois): routing + POI attribute questions. Great regression
suite for the QA sub-feature.
- **TravelPlanner** (ICML'24, OSU-NLP-Group/TravelPlanner) — 1,225 curated
travel-planning intents against ~4M records with a tool sandbox; the
standard benchmark for exactly this task class (constrained, multi-day,
transportation + meals + attractions + lodging). Use its constraint-checker
idea for our itinerary validator.
- Build a small local eval set over our region with golden itineraries;
score: (a) tool-answer correctness (deterministic, easy), (b) validator
pass-rate of LLM-edited itineraries (should be high after validation),
(c) LLM faithfulness to tool results (LLM-as-judge + spot checks),
(d) refinement quality: does a 1-instruction edit touch only the intended
stops (diff-locality metric).
- Track: question types, tool-call success rate, hallucinated-number rate
(should be ~0), median latency.
## Model selection (server-first, on-device as fallback)
Server (primary): any strong function-calling model behind one swappable
interface. Candidates from current benchmarks:
- **Qwen3.8-27B** (dense, ~24 GB) — **first choice for this system: it runs
well on our hardware.** The 27B class buys real headroom exactly where
this workload stresses models: long context (whole-trip summaries across
places/days), multi-step planning with many tool rounds, and taste-level
suggestion writing. Good general/search reasoning; native tool calling.
- **Granite 4.2-8B** (Aug 2026): dense, Apache 2.0, native OpenAI-format
tool calling, agentic RL training, switchable thinking. Strong fit for the
refine-loop workload at lower compute; run at Q4 or FP16, thinking off on
simple tool turns.
- **Qwen3-30B-A3B-Instruct-2507**: MoE with 3B active — 8B-class capacity at
3B-class latency; great if we serve on modest GPUs.
- Any hosted API (Claude/GPT-class) as the "quality ceiling" option and
for LLM-as-judge in eval.
On-device fallback (optional, later — e.g. offline trip planning on a phone):
- **Qwen3-4B-Instruct-2507** (dense 4B, Q4 ≈ 2.53 GB) — the reference
small tool-calling model; note it is the *4B*, not a 3B (the "3B" in the
family is the 30B-A3B MoE's *active* count).
- **Granite 4.2-3B** — new contender with the same agentic-RL pedigree; worth
a head-to-head on our MapQA subset (don't trust vendor BFCL numbers —
Granite 4.1-8B's were panned by the community).
- **Phi-4-mini (3.8B)** — fastest of the class, best if latency dominates.
- Sub-2B models: not reliable agent brains (tool JSON breaks); acceptable
only as an intent classifier in front of the main model.
- Quantization floor: **never below 4-bit** — structured output (tool calls)
degrades fastest under aggressive quantization (QuantCall benchmark; Q3
already shows 38% PPL + visible reasoning regression), and a 23-bit 8B
is strictly worse than a 4B at Q4 for the same footprint.
Reality check that shaped the server-first choice: an on-device 4B agent
loop takes ~3060 s per multi-hop question (each tool call is a 100300
token generation at 815 tok/s). Fine as an offline fallback for "is this
open?"-class questions; the itinerary planner belongs on the server.
## Build plan
**Pre-v0 (mock, 12 days):** before backend work, a **static interactive
mock of the web UI** in the real frontend stack, with fake data behind the
same interface the API will implement: map + timeline rail + chat; vibe
chips; the **landing screen** ("Where would you like to go?" + free-form field
+ drop zone, incl. a mocked booking drag-in → parse → anchored plan);
suggestion cards with *simulated streaming enrichment* (image +
citation appearing after a delay); accept/reject diff animation;
like/rule-out reactions; **drag-reorder a stop and watch leg durations +
the day budget bar update live**; **L3 route editor (hike: draggable
trail with live duration; urban: multi-modal walk+bus leg)**;
**drop a (mocked) train ticket → parsed transit object pinned, day
re-anchors around it**; a stop card showing planned vs suggested time +
price; **issues panel with one mocked warning ("lunch 45 min, typically
~90 at this venue") and a one-click fix**; L0/L1 drill-down (mocked
data only). Purpose:
validate the interaction model, settle the visual design language, and
produce the API data contract by construction. The mock becomes the app shell — only the fake data layer
gets swapped for the real server.
**v0 (days, all-hosted APIs):** chat + Leaflet + day timeline; the
**hierarchical skeleton from day one** (trip → place → days → stops,
start with one place); exploration loop: destination → OSM tag stats →
vibe chips → suggestion cards → `preferences`; single-day itinerary over
Nominatim (demo) + Overpass + GraphHopper public API + opening-hours lib;
QA mode + add/move/remove stop + validator. Validation: the fluid
"where would you like to go → vibe → place plan" flow works end-to-end for
one city.
**v1 (local data + multi-place):** PostGIS + local GraphHopper + local
Nominatim for the target region; multi-day + **multi-place itineraries with
opaque `transit`** and junction validation; `plan_day` heuristic planner,
versioning/undo, semantic POI index (pgvector); TravelPlanner-style
constraint eval in CI.
**v2 (multi-country + mobile on-trip):** per-place data loading for true
multi-country trips; external rail/air transit APIs; the mobile app with
today-view, leg navigation, **offline trip bundle** (tiles + itinerary +
stop data + hours), and offline-edit sync; alternatives UI (structured
option cards) in web; clarifying questions; "open on arrival" filters; MapQA
+ local eval in CI; optional VLM: send the *rendered* map view to a VLM
only when the user shares/asks about a screenshot; optional on-device 4B
for offline "is X open?".
## Tech stack suggestion (aligns with your existing Go tooling)
- **Go** service: agent loop, tools, HTTP API (chi), Postgres driver with
PostGIS, GraphHopper/OSRM client, `osm-opening-hours` equivalent
(or shelling to the Python lib / a small sidecar).
- **LLM**: any function-calling model; keep it behind one interface so the
provider is swappable (including a local model for offline use).
- **Frontend**: small SPA (or Go + templ + htmx if you want minimal) with
Leaflet.
## References
- MapQA — arxiv.org/abs/2503.07871
- RAG architectures for POI QA — talks.osgeo.org/foss4g-2026/talk/VBGM3W
- MapBench (VLM map reading limits) — arxiv.org/abs/2503.14607
- MapReason-OSM — arxiv.org/abs/2606.22597
- LAMP: A Language Model on the Map — arxiv.org/abs/2403.09059
- OSRM API — project-osrm.org/docs
- Overpass API — wiki.openstreetmap.org/wiki/Overpass_API
- GraphHopper — github.com/graphhopper/graph-hopper
- GROKE (OSM graph reasoning, ACL'26) — github.com/Geo-R2LLM/groke
- TravelPlanner (ICML'24 travel-planning benchmark) — github.com/OSU-NLP-Group/TravelPlanner
- Granite 4.2 (IBM open reasoning models, Aug 2026) — ibm.com/granite/docs/models/granite4-2
- Qwen3-2507 refresh (Instruct/Thinking, 4B235B) — github.com/QwenLM/Qwen3
- TravelAgent (LLM travel planning, 2024) — jiangjiechen.github.io/publication/travelagent
- QuantCall (quantization effect on tool calling) — github.com/Happynood/quant-toolcall-bench