Palate — a personal wine operating system
Design spec · 2026-07-09 · status: awaiting user approval
Working name "Palate" is a placeholder; naming comes later.
1. Philosophy
One-sentence definition: a sommelier-analyst you talk to like a ChatGPT thread — except every conversation deposits structured, permanent data about your palate, and the AI's opinions about you are backed by that record instead of vibes.
The inversion: every existing wine app treats the wine as the record and your opinion as metadata. Palate treats the tasting as the record and the wine as metadata. The chat is the UI; the instrument is the memory.
Origin evidence: a real ChatGPT thread ("Best Value Wine Picks") proved the interaction loop works from message one — report a wine + context → AI infers palate → explains why → converts to tiered, priced buy/skip calls against a real retailer. What the thread could not do: persist its palate model, measure anything, or resist flattering. Palate keeps the conversational magic and fixes those three failures.
Three commitments
- Talk first, structure always. Free text, voice, or a label photo is a complete log. The AI converts it to SAT-structured data invisibly, asking at most a couple of sharpening questions. Deliberate SAT-grid work is a mode you choose (exam training), never the toll for entry.
- Opinions must cite the record. Every AI claim about the taster links to the tastings that support it — falsifiable, unlike a chat thread. Blind mode grades conclusions against ground truth and shows the miss. Honesty over encouragement.
- The AI is both concierge and analyst. Live: diagnose a wine, mine a retailer list / restaurant list / auction sheet for value, tiered buy-skip with prices. Overnight: the deeper analyst pass — calibration drift, palate patterns, WSET readiness. Runs on plan-auth Claude ($0 marginal); live chat tolerates occasional rate-limit pauses.
What it is not (permanent non-goals)
Not social (no feeds, no followers). Not inventory software (a <30-bottle shelf gets a list, nothing more). Not a browsable encyclopedia (knowledge attaches to wines tasted and units studied). Not a chat thread that forgets. Single-taster instrument: companions are modeled, never onboarded.
Success criterion
In twelve months: blind-tasting calibration error measurably down, L3 booked on evidence — and logging never once felt like homework.
2. User reality (design inputs)
- WSET Level 2 in progress, exam within weeks; L3 then Diploma intended.
- Existing habit: unstructured notes — the product structures a habit that already exists.
- Shelf of <30 bottles; buy-to-drink value hunter (Costco/Kirkland, K&L incl. auctions; post-move: LCBO Vintages + Ontario agencies).
- Relocating California → Toronto imminently; retailer-aware recommendations must be LCBO-first at launch.
- Parent of infant twins: time-poor; capture must survive a ten-second dinner-table window.
- Palate (seed priors for the model, to be confirmed by data): elegant / silky / high-acid / finesse over power; grower Champagne, precise white Burgundy, Loire Chenin; dislikes heavy oak and stem-forward "bark and twigs" styles. Burgundy sweet spot $50–100.
- Wife is the primary drinking companion; her reactions are worth modeling.
- Form factor: phone-first self-hosted PWA (capture) + desktop view (study/review).
- AI budget: plan-auth Claude, $0 marginal; architecture must be async-tolerant.
3. Core object model
Seven objects carry the product. Everything else is a view.
- Wine — identity of a bottling: producer, cuvée, appellation, vintage, grape(s), price paid, retailer. Deliberately thin; exists so multiple tastings link up.
- Tasting — the atomic record. Stores two representations permanently: (a) the raw input verbatim — text, voice transcript, label photo — never discarded; (b) the SAT structure (appearance/nose/palate/conclusions in WSET lexicon), each field tagged with who asserted it (user deliberately vs. AI-inferred) and at which SAT level. Plus context (place, food, companions, pairing success/fail tag) and mode:
capture (AI structures your words), practice (you fill the grid), blind (practice + hidden identity + reveal). Fast negative capture ("ugh, skip" in ten seconds; AI diagnoses later) is a first-class capture path. A Tasting carries a reactions list: per person, verdict + verbatim comment, extracted from natural phrasing.
- Calibration record — generated only by deliberate work: practice grids (scored by AI critique), blind grids (scored vs. the revealed bottle), and pre-pour predictions (scored vs. the subsequent log). Includes per-attribute error over time and predicted-confidence vs. accuracy (catches overconfidence). Casual captures never generate calibration data — you can't be graded on words the AI structured for you.
- Palate model — persistent, versioned profile of discrete claims, each with confidence and links to supporting tastings ("prefers energy/finesse over power — evidence: 14 tastings"). Updated by the overnight analyst; cited by the live concierge.
- Study state — WSET level, exam date, syllabus units, spaced-repetition queue. Items are preferentially — not exclusively — generated from tasted wines; pure theory (classification systems, viticulture) gets cards no bottle produces. Includes syllabus-coverage map (studied vs. tasted) and retention scores.
- Shelf & wishlist — on-hand list including open-bottle state with drink-by nudge; wishlist entries carry the concierge's reasoning. Buy queue is steered by value and syllabus coverage. Haul capture is the bulk entry path: one photo of an entire wine haul (bottles lined up, or the receipt) → AI identifies every bottle → user confirms/corrects the list in one pass → all Wine records created and added to the Shelf with price/retailer where legible. Doubles as onboarding: seeding the new Toronto shelf is one photo, not twenty forms.
- People & companion profiles — small roster of recurring companions (wife first-class). Companion profile = claims-with-evidence at lower resolution, built only from reported reactions; no SAT, no calibration, no study state. Profiles state their epistemic basis (second-hand) and never claim self-report confidence. Enables intersection queries ("what will we both love") and divergence flags.
The loop: raw input → Tasting → (practice/blind/prediction) → Calibration → Palate model → concierge recommendations → Shelf/wishlist → next bottle → raw input. One loop, compounding.
SAT level is a dial, not a gate
v1 ships the full L3-grade SAT schema (WSET levels nest; L2 is a coarse subset). Capture structures to the deepest level the input supports. Practice mode has a rigor dial (L2/L3); calibration is scored per level, so the L3-readiness trendline starts accumulating from day one. Honesty guard: the analyst calls out when L3 tasting practice is crowding out nearer-term L2 theory gaps.
The verification ladder
Every graded item carries a truth_source rung, displayed in the UI:
- Rung 1 — your own data (proven). Own-history recall vs. stored records; predictions vs. subsequent log (openly labeled self-consistency); blind identity calls vs. the revealed bottle. The record is the judge; the AI only compares and explains.
- Rung 2 — the syllabus (keyed). Theory answers grade against a curated fact bank built from official WSET materials, never against the AI's general knowledge (hallucinated grades are worse than none). AI reads free-text answers vs. the key and explains the why. Grades are contestable in chat; contested items are flagged and re-checked with sources.
- Rung 3 — no oracle (refereed). Open-tasting SAT assessments are checked against typicity priors, documented consensus where it exists, and the user's own track record — returned as "plausible/unusual, here's why" with confidence and sources shown, never a red X. Structure (acid/tannin/sweetness/body) gets firm refereeing; aroma descriptors the lightest touch — "I smell violets" is not falsifiable, and teaching distrust of one's own nose is anti-goal.
Knowing whether you were proven right, keyed right, or plausibly right is itself WSET training: L3 grades defensible conclusions from evidence.
4. Experience pillars — five surfaces
Phone-first PWA; the conversation is the home screen. Everything is reachable from chat; surfaces exist where a grid, graph, or list beats prose. Design rule: other surfaces are faster, never required.
- The Conversation — front door and only mandatory surface. Capture, concierge (label diagnosis; "what should I grab at LCBO"; restaurant wine-list photo → pick in <30s using palate model + companion profile + food + budget; retailer list/PDF/auction mining; haul photo → bulk shelf add; occasion modes), questions, contesting grades.
- The Practice Room — deliberate mode. Choose rigor (L2/L3) and mode (open/blind), optional exam timer; you fill the grid, AI silent until commit, then critiques content and form and walks the climate→structure→variety elimination logic; blind reveals and scores. Calibration history lives here: per-attribute error, confidence-vs-accuracy, prediction-vs-actual, the L3-readiness trendline.
- The Shelf — on hand (incl. open bottles + drink-by), wishlist with reasoning, buy queue steered by value + syllabus. "What should I open tonight with [wife]" resolves here.
- The Desk — syllabus map with coverage (studied ∕ tasted), spaced-repetition queue, the Daily Sip, exam countdown with the analyst's honest readiness read, retention scores.
- The File — the palate model made visible: claims with evidence trails, companion profiles, archive of Analyst's Letters. The surface that proves "it remembers the taster."
Testing & retention (spans Desk + Practice Room)
Retrieval practice, spaced not periodic, attached to moments that already exist:
- Daily Sip — one 60-second card daily (theory item due, or own-history question, alternating). Skippable without guilt; streak-free by design.
- Own-history questions — retrieval over the personal record ("what grape was the Protégé wine?").
- Predict before you pour — 15-second structural prediction when opening a Shelf bottle, scored against the subsequent log; turns casual bottles into micro blind-tests with zero ceremony. The mechanic this design defends hardest.
- Weekly Letter retention check — three questions on that week's learnings.
- Anti-homework guard: testing never manufactures new demands on attention; triggers are real moments (a bottle opened, a morning coffee, a Sunday letter).
The Analyst's Letter
Weekly push, archived in The File: what your palate did this week, calibration drift, retention read, one thing to taste next and why, L2-theory honesty check, monthly spend/value digest.
5. v1 scope and build order
Per user decision (2026-07-09), the former Phase 2 is merged into v1. The L2 exam date is protected by sequencing, not scope-cutting: exam-critical tranches ship first.
Tranche A — before the L2 exam (exam-critical, ships first):
- Capture, complete: chat/voice/label-photo → SAT structuring (full L3 schema), reactions extraction, fast negative capture.
- The Desk, L2-focused: curated L2 fact bank from official materials, spaced repetition, Daily Sip, exam countdown, readiness reads.
- Concierge core: diagnose, tiered buy/skip, list mining incl. restaurant wine-list photo; LCBO-aware.
- Palate model v1 + weekly Analyst's Letter (overnight, plan-auth).
- Practice minimum honest dose: open SAT practice with critique, prediction-before-pour, basic blind (hide → grade vs. reveal). Verification-ladder tagging from day one (retrofitting epistemics never works).
- Shelf lite: on-hand, open-bottle state, wishlist with reasoning — seeded via haul capture (photo of a full haul or receipt → confirm → shelf populated). Launch-critical: an empty shelf weakens every loop downstream, and the Toronto restock is the natural seeding moment.
Tranche B — immediately after (same v1, lands behind the exam):
- Calibration dashboards (per-attribute error, confidence-vs-accuracy, L3-readiness trendline).
- Timed exam mode + elimination-scaffold coaching + answer-form feedback.
- Syllabus-coverage steering of the wishlist ("Vintages has a textbook Fino for $16 — you've never tasted flor").
- Occasion parameter on the concierge (gift / crowd / anniversary — flips out of value-hunting).
- Companion intersection queries + divergence flags.
v1 success criteria: L2 passed; ≥40 tastings logged; capture never skipped because it felt like work; Tranche B live within weeks of the exam.
Acknowledged risk: merged v1 is roughly double the original cut. Mitigation is the tranche order above; if the exam date pressures the build, Tranche B slips, never Tranche A.
6. Roadmap — mapped to the WSET campaign
v1 (now → weeks after L2): everything in §5.
Phase L3 campaign (enrollment → tasting exam):
- L3 rigor as default; mock tasting exams graded against the WSET rubric with answer-form coaching.
- Aroma trainer (descriptor-to-smell drilling, paired with a physical aroma kit).
- Comparative flight builder (plan a bottle set isolating one variable; diff SAT grids side by side).
- Readiness gate: the analyst says when blind-conclusion hit-rate supports booking the exam.
- Success: L3 passed with Merit or better; the tasting paper feels familiar.
Phase Diploma era (multi-year):
- D1–D6 theory packs with essay-form coaching (Diploma grades arguments, not recall).
- Long-horizon palate analytics: vintage variation personally tasted; multi-year palate drift.
- The instrument co-authors its roadmap from its own data.
- Success: still in daily use in year three — the compounding bet paid.
7. Parking lot (rejected, with revisit triggers)
| Item |
Why rejected |
Revisit trigger |
| Wine-fault drills |
Can't ship smells; reference content drifts encyclopedia-ward |
First corked/faulty bottle logged → teachable-moment capture; aroma-kit era |
| Drinking-window tracking |
Inventory-software territory at <30 bottles |
Shelf regularly >75 bottles or age-worthy buying becomes a pattern |
| Travel/winery mode |
Capture already works anywhere |
Regular winery trips resume (Niagara/PEC/CA) |
| Full cellar management |
Shelf list suffices at current scale |
Same trigger as drink windows |
8. Design quality bar
The founding brief's standard — "feels like Apple designed it": opinionated, minimal, coherent; a product that feels inevitable rather than assembled. Commitments the build must honor:
- Designed through Claude Design (claude.ai/design), Claude's native design tooling — the app's design system lives as a Claude Design project (tokens, type scale, components as reviewable cards), kept in sync with the app's component library via the native design-sync workflow, component by component. Not ad-hoc styling; not the operator's personal skillstack (explicit user decision, 2026-07-09). Design direction still gets its own short brief with reference anchors gathered early, before UI build starts.
- Typography baseline, always-on: no widows/orphans, curly quotes, real em/en dashes, tabular numerals where data aligns.
- Accessibility & motion hard gate:
prefers-reduced-motion honored; no content gated behind motion reveals; the reduced state is designed, not gutted. Capture flow usable one-handed, at a dinner table, in low light (dark mode is a first-class state, not an afterthought — wine is drunk at night).
- Calm instrument, not a dashboard: the aesthetic serves the ten-second capture and the quiet weekly letter; no gamification chrome (no streaks, badges, confetti).
- Visual verification before any "done" claim: screenshot-verified on real phone viewport sizes; design judged blind against the bar, never self-certified.
9. Explicitly out of scope for this document
Tech stack, architecture, hosting, schema DDL, model/prompt design — all deferred until this vision is approved (per the founding brief: no code, no scaffolding, no stack decisions).