Role: Main developer in a team of four senior engineers — designed and built the core platform end to end: UI and frontend, the NestJS + FastAPI backend, the LangGraph AI service, data, and CI/CD.
Next.js application with SSG/ISR for 95+ Lighthouse performance, deployed on Hetzner + Cloudflare Workers. A real-time SSE streaming UI shows meals being generated live; responsive design uses shadcn/ui components with smooth animations.
AI Generation Pipeline
One tool-using LangGraph agent, traced end-to-end — the request flows top to bottom through nine layers, wrapped by two cross-cutting rails.
Client / API Gateway
Receives the request, enqueues a job, and POSTs to the Python service.
Interface / Serving
SSE endpoint that streams generation events back, behind a configurable 360 s hang watchdog.
Orchestration / Control-flow
A LangGraph StateGraph drives the day loop and the validate-retry loop.
Reasoning / Model
The ReAct decision loop — reason, act, observe, repeat.
Tooling / Action
Forced tool use to fetch ground-truth nutrition before committing a day.
The LLM never writes nutrition numbers — the server computes them from USDA ground truth.
A hard architectural constraint: the submit schema has no nutrition fields, so the model cannot fabricate macros — stronger than any prompt instruction.
Knowledge / Retrieval (RAG)
Two complementary RAG systems — exact ingredient macros vs open-ended dish ideas.
Exact ingredient → macro grounding.
Open-ended dish ideas before generation.
two RAGs, two query shapes — exact ingredient macros vs open-ended dish ideas
Validated meals are re-ingested into the recipe corpus via a LlamaIndex pipeline — future generations retrieve from real past outputs, not just the seed corpus.
Drift guards prevent mode collapse: provenance weighting (×0.7), a ≤1-of-3 generated cap, semantic dedup (cosine >0.95), and a weekly variety audit with clean rollback.
Memory / State / Persistence
Typed graph state checkpointed to Postgres — an interrupted plan resumes from the next unfinished day.
Guardrails / Validation
Schema enforcement plus dietary-restriction checks on every output.
Evaluation / Quality
A 3-layer online validate-and-retry cascade, backed by a pre-release eval harness.
9. Observability
cross-cutting · spans every layerTraces every LLM call across the whole stack.
10. Data / Infrastructure
cross-cutting · spans every layerRuns on Hetzner + Cloudflare Workers (moved from Railway in August 2026).
AI System
The graph, tools and retrieval behind a meal plan. A LangGraph StateGraph generates the plan one day at a time, a ReAct tool loop grounds each day in USDA data, and two retrieval stores answer two different questions: which food is this, and what should we cook.
Meal-generation graph
LangGraph StateGraph · meal_generation.pyBuilds the per-plan context in graph state before the first day.
Joins the graph when recipe RAG is on: pulls candidate dishes from Qdrant into state.
One day of meals, produced by a tool-using ReAct loop.
The loop ends when the model calls the forced submit tool DayPlanLLMSubmit (bind_tools(…, tool_choice=…)). Its schema has no nutrition fields.
3-layer check. On failure, a conditional edge sends structured feedback back to generate_day; the oscillation detector stops a day that flips between the same failures.
Streams the finished day out as an SSE event.
Records the day, then a conditional edge loops to the next day or ends the run.
Each participant is generated in parallel, then the results are merged and reconciled into one household plan (4+ participants).
Tools
- lookup_usda_batch — Many ingredients in one call; schema-enforced batch input.
- search_usda_food — Search USDA foods by name.
- get_usda_nutrition — Nutrition values for one USDA food.
- DayPlanLLMSubmit — Forced submit tool that ends the day. No nutrition fields.
A separate LangChain / LangGraph agent that answers questions and acts on the user's plans, e.g.:
RAG over two stores
Its own PostgreSQL 17. Maps an ingredient string to the right USDA food, so every macro traces back to USDA data.
- full-text search: to_tsvector · plainto_tsquery · ts_rank
- top 40 candidates
- Claude LLM reranker picks the match
Qdrant over ~230k filtered RecipeNLG recipes. Feeds dish ideas into retrieve_recipes; it shapes dish choice, never the numbers.
- tag pre-filter
- ANN search per collection · 2 collections
- RRF fusion across both result lists
- post-filter
- LLM rerank
State and serving
Key Features
- Next.js App Router with SSG/ISR achieving 95+ Lighthouse performance
- Real-time SSE streaming UI showing meals generated live with progressive rendering
- Responsive design with shadcn/ui component library
- Smooth animations and transitions for meal plan interactions
- SEO-optimized pages with structured data and meta tags
Tech Stack
Frontend
Infrastructure
Backend
AI Orchestration
RAG & Retrieval
Eval & Observability
Challenges & Solutions
Real-Time Streaming Across Services
Meal generation takes 10-30s through the AI pipeline. Users need immediate feedback, but the response crosses 3 service boundaries (Python → NestJS → Next.js).
SSE streaming pipeline where the FastAPI ai-service emits Server-Sent Events to NestJS, which proxies them to the Next.js frontend behind a configurable 360 s hang watchdog. Users see meals being generated in real-time.
Staying Up When the LLM Provider Fails
A production outage when the Anthropic account balance ran out showed that every generation depended on a single LLM provider.
Added a live cross-provider fallback — Anthropic ⇄ a self-hosted model — so generation switches providers instead of failing. It has fired in production when the Anthropic balance ran out.
Two Retrieval Jobs: Nutrition vs Dish Ideas
Exact ingredient→macro lookup and open-ended dish inspiration are different retrieval problems. Short ingredient names ('chicken breast, raw') are won by lexical search, while conceptual dish queries ('Moroccan chickpea stew' ≈ 'North African lentil soup') are lexically distant but semantically equivalent, so full-text search misses them — one retrieval mechanism can't serve both.
Two complementary retrieval paths that don't compete. The factual layer is USDA lookup — Postgres full-text search plus an LLM rerank — grounding every macro. The creative layer is recipe retrieval on Qdrant over ~230k filtered RecipeNLG recipes, fusing two collections (~229k and ~2.1M points), whose results feed dish ideas into generation. USDA still grounds every number; recipe retrieval only shapes dish selection.
Reliable AI Meal Generation
An LLM is non-deterministic, yet a multi-day meal plan needs a reliable, repeatable pipeline with tool use, retries, and streaming — a hand-rolled while-loop gets brittle fast.
Modeled generation as a LangGraph StateGraph: a deterministic day loop (prepare_context → generate_day ⇄ usda_tools → validate_day → emit_day → update_history) wraps a ReAct-style inner loop where the LLM decides tool calls, and a forced submit tool cleanly terminates each day.
Preventing Hallucinated Nutrition Numbers
LLMs confidently fabricate plausible-but-wrong calorie and macro values, and a meal planner that lies about nutrition is worse than useless.
The LLM's submit schema (DayPlanLLMSubmit) has no nutrition fields, so the model physically can't write a macro number — it must call a USDA lookup tool for ground truth and the server computes every calorie and macro. A hard architectural constraint, not a prompt instruction.
Enforcing Quality on Non-Deterministic Output
A generated day can still miss calorie/macro targets or violate dietary restrictions (keto, vegan, allergen-free), and a single LLM pass can't be trusted.
A 3-layer validate-and-retry cascade — programmatic checks → optional LLM-as-judge (Haiku) → a blocking USDA ground-truth fact-check that retries at >50% calorie or >60% macro deviation — re-prompts with structured feedback injected into the next attempt, with an oscillation detector to bail on stuck days.
Crash-Resumable Long-Running Generation
Generating a full multi-day plan is long-running, and a crash or timeout mid-plan would otherwise discard every day already produced.
Graph state is checkpointed to Postgres via AsyncPostgresSaver, keyed by thread_id, so an interrupted plan resumes by restarting from the next unfinished day instead of starting over.
Preventing Mode Collapse in the Data Flywheel
Re-ingesting the model's own validated meals into the recipe corpus closes a data flywheel, but conditioning future generation on past generations risks an ever-narrowing dish distribution — mode collapse, where the system retrieves and re-generates the same handful of dishes in a tightening loop.
A guarded self-ingestion ETL: only meals that pass all 3 cascade layers with <5% deviation, no oscillation, and no force-commit qualify. Drift guards prevent collapse — provenance weighting (generated recipes ×0.7 at retrieval), a fraction cap (≤1 of 3 injected recipes may be generated), semantic dedup (skip if cosine >0.95), and a weekly variety audit with a provenance audit trail for clean rollback. Feature-flagged until 30-day variety metrics prove stable.