Palate — a personal wine operating system

Design spec · 2026-07-09 · status: awaiting user approval Working name "Palate" is a placeholder; naming comes later.

1. Philosophy

One-sentence definition: a sommelier-analyst you talk to like a ChatGPT thread — except every conversation deposits structured, permanent data about your palate, and the AI's opinions about you are backed by that record instead of vibes.

The inversion: every existing wine app treats the wine as the record and your opinion as metadata. Palate treats the tasting as the record and the wine as metadata. The chat is the UI; the instrument is the memory.

Origin evidence: a real ChatGPT thread ("Best Value Wine Picks") proved the interaction loop works from message one — report a wine + context → AI infers palate → explains why → converts to tiered, priced buy/skip calls against a real retailer. What the thread could not do: persist its palate model, measure anything, or resist flattering. Palate keeps the conversational magic and fixes those three failures.

Three commitments

  1. Talk first, structure always. Free text, voice, or a label photo is a complete log. The AI converts it to SAT-structured data invisibly, asking at most a couple of sharpening questions. Deliberate SAT-grid work is a mode you choose (exam training), never the toll for entry.
  2. Opinions must cite the record. Every AI claim about the taster links to the tastings that support it — falsifiable, unlike a chat thread. Blind mode grades conclusions against ground truth and shows the miss. Honesty over encouragement.
  3. The AI is both concierge and analyst. Live: diagnose a wine, mine a retailer list / restaurant list / auction sheet for value, tiered buy-skip with prices. Overnight: the deeper analyst pass — calibration drift, palate patterns, WSET readiness. Runs on plan-auth Claude ($0 marginal); live chat tolerates occasional rate-limit pauses.

What it is not (permanent non-goals)

Not social (no feeds, no followers). Not inventory software (a <30-bottle shelf gets a list, nothing more). Not a browsable encyclopedia (knowledge attaches to wines tasted and units studied). Not a chat thread that forgets. Single-taster instrument: companions are modeled, never onboarded.

Success criterion

In twelve months: blind-tasting calibration error measurably down, L3 booked on evidence — and logging never once felt like homework.

2. User reality (design inputs)

3. Core object model

Seven objects carry the product. Everything else is a view.

  1. Wine — identity of a bottling: producer, cuvée, appellation, vintage, grape(s), price paid, retailer. Deliberately thin; exists so multiple tastings link up.
  2. Tasting — the atomic record. Stores two representations permanently: (a) the raw input verbatim — text, voice transcript, label photo — never discarded; (b) the SAT structure (appearance/nose/palate/conclusions in WSET lexicon), each field tagged with who asserted it (user deliberately vs. AI-inferred) and at which SAT level. Plus context (place, food, companions, pairing success/fail tag) and mode: capture (AI structures your words), practice (you fill the grid), blind (practice + hidden identity + reveal). Fast negative capture ("ugh, skip" in ten seconds; AI diagnoses later) is a first-class capture path. A Tasting carries a reactions list: per person, verdict + verbatim comment, extracted from natural phrasing.
  3. Calibration record — generated only by deliberate work: practice grids (scored by AI critique), blind grids (scored vs. the revealed bottle), and pre-pour predictions (scored vs. the subsequent log). Includes per-attribute error over time and predicted-confidence vs. accuracy (catches overconfidence). Casual captures never generate calibration data — you can't be graded on words the AI structured for you.
  4. Palate model — persistent, versioned profile of discrete claims, each with confidence and links to supporting tastings ("prefers energy/finesse over power — evidence: 14 tastings"). Updated by the overnight analyst; cited by the live concierge.
  5. Study state — WSET level, exam date, syllabus units, spaced-repetition queue. Items are preferentially — not exclusively — generated from tasted wines; pure theory (classification systems, viticulture) gets cards no bottle produces. Includes syllabus-coverage map (studied vs. tasted) and retention scores.
  6. Shelf & wishlist — on-hand list including open-bottle state with drink-by nudge; wishlist entries carry the concierge's reasoning. Buy queue is steered by value and syllabus coverage. Haul capture is the bulk entry path: one photo of an entire wine haul (bottles lined up, or the receipt) → AI identifies every bottle → user confirms/corrects the list in one pass → all Wine records created and added to the Shelf with price/retailer where legible. Doubles as onboarding: seeding the new Toronto shelf is one photo, not twenty forms.
  7. People & companion profiles — small roster of recurring companions (wife first-class). Companion profile = claims-with-evidence at lower resolution, built only from reported reactions; no SAT, no calibration, no study state. Profiles state their epistemic basis (second-hand) and never claim self-report confidence. Enables intersection queries ("what will we both love") and divergence flags.

The loop: raw input → Tasting → (practice/blind/prediction) → Calibration → Palate model → concierge recommendations → Shelf/wishlist → next bottle → raw input. One loop, compounding.

SAT level is a dial, not a gate

v1 ships the full L3-grade SAT schema (WSET levels nest; L2 is a coarse subset). Capture structures to the deepest level the input supports. Practice mode has a rigor dial (L2/L3); calibration is scored per level, so the L3-readiness trendline starts accumulating from day one. Honesty guard: the analyst calls out when L3 tasting practice is crowding out nearer-term L2 theory gaps.

The verification ladder

Every graded item carries a truth_source rung, displayed in the UI:

Knowing whether you were proven right, keyed right, or plausibly right is itself WSET training: L3 grades defensible conclusions from evidence.

4. Experience pillars — five surfaces

Phone-first PWA; the conversation is the home screen. Everything is reachable from chat; surfaces exist where a grid, graph, or list beats prose. Design rule: other surfaces are faster, never required.

  1. The Conversation — front door and only mandatory surface. Capture, concierge (label diagnosis; "what should I grab at LCBO"; restaurant wine-list photo → pick in <30s using palate model + companion profile + food + budget; retailer list/PDF/auction mining; haul photo → bulk shelf add; occasion modes), questions, contesting grades.
  2. The Practice Room — deliberate mode. Choose rigor (L2/L3) and mode (open/blind), optional exam timer; you fill the grid, AI silent until commit, then critiques content and form and walks the climate→structure→variety elimination logic; blind reveals and scores. Calibration history lives here: per-attribute error, confidence-vs-accuracy, prediction-vs-actual, the L3-readiness trendline.
  3. The Shelf — on hand (incl. open bottles + drink-by), wishlist with reasoning, buy queue steered by value + syllabus. "What should I open tonight with [wife]" resolves here.
  4. The Desk — syllabus map with coverage (studied ∕ tasted), spaced-repetition queue, the Daily Sip, exam countdown with the analyst's honest readiness read, retention scores.
  5. The File — the palate model made visible: claims with evidence trails, companion profiles, archive of Analyst's Letters. The surface that proves "it remembers the taster."

Testing & retention (spans Desk + Practice Room)

Retrieval practice, spaced not periodic, attached to moments that already exist:

The Analyst's Letter

Weekly push, archived in The File: what your palate did this week, calibration drift, retention read, one thing to taste next and why, L2-theory honesty check, monthly spend/value digest.

5. v1 scope and build order

Per user decision (2026-07-09), the former Phase 2 is merged into v1. The L2 exam date is protected by sequencing, not scope-cutting: exam-critical tranches ship first.

Tranche A — before the L2 exam (exam-critical, ships first):

Tranche B — immediately after (same v1, lands behind the exam):

v1 success criteria: L2 passed; ≥40 tastings logged; capture never skipped because it felt like work; Tranche B live within weeks of the exam.

Acknowledged risk: merged v1 is roughly double the original cut. Mitigation is the tranche order above; if the exam date pressures the build, Tranche B slips, never Tranche A.

6. Roadmap — mapped to the WSET campaign

v1 (now → weeks after L2): everything in §5.

Phase L3 campaign (enrollment → tasting exam):

Phase Diploma era (multi-year):

7. Parking lot (rejected, with revisit triggers)

Item Why rejected Revisit trigger
Wine-fault drills Can't ship smells; reference content drifts encyclopedia-ward First corked/faulty bottle logged → teachable-moment capture; aroma-kit era
Drinking-window tracking Inventory-software territory at <30 bottles Shelf regularly >75 bottles or age-worthy buying becomes a pattern
Travel/winery mode Capture already works anywhere Regular winery trips resume (Niagara/PEC/CA)
Full cellar management Shelf list suffices at current scale Same trigger as drink windows

8. Design quality bar

The founding brief's standard — "feels like Apple designed it": opinionated, minimal, coherent; a product that feels inevitable rather than assembled. Commitments the build must honor:

9. Explicitly out of scope for this document

Tech stack, architecture, hosting, schema DDL, model/prompt design — all deferred until this vision is approved (per the founding brief: no code, no scaffolding, no stack decisions).