Back to Projects

Melio MealPlan AI

AI-Powered Meal Planning

Overview

Role: Main developer in a team of four senior engineers — designed and built the core platform end to end: UI and frontend, the NestJS + FastAPI backend, the LangGraph AI service, data, and CI/CD.

What it is. Full-stack AI meal-planning platform built as a pnpm monorepo and running on Hetzner + Cloudflare Workers — 311 generations from 105 users by September 2026.

Architecture. NestJS handles business logic, a Python FastAPI ai-service runs the LangGraph agent with server-computed USDA nutrition and Qdrant recipe retrieval, and Next.js delivers a performant frontend with real-time SSE streaming.

Serving. NestJS is the API and queue layer: each plan is a BullMQ job handed to the Python FastAPI AI service (LangGraph), which streams finished days back over SSE. In August 2026 the platform moved from Railway to Hetzner + Cloudflare Workers and gained a cross-provider LLM fallback.

AI Generation Pipeline

One tool-using LangGraph agent, traced end-to-end — the request flows top to bottom through nine layers, wrapped by two cross-cutting rails.

NestJS API → BullMQ job → Python AI service
0

Client / API Gateway

Receives the request, enqueues a job, and POSTs to the Python service.

NestJSBullMQ + Redis
1

Interface / Serving

SSE endpoint that streams generation events back, behind a configurable 360 s hang watchdog.

FastAPIasyncioSSE
2

Orchestration / Control-flow

A LangGraph StateGraph drives the day loop and the validate-retry loop.

LangGraph StateGraphconditional edgesreducers
prepare_contextretrieve_recipesgenerate_dayReActusda_toolsvalidate_dayemit_dayupdate_history
validate_day retries back to generate_day with structured feedback
scales out via supervisor-workers (LangGraph Send()) for 4+ participant households
3

Reasoning / Model

The ReAct decision loop — reason, act, observe, repeat.

ChatAnthropicClaude Haiku 4.5 / Sonnetself-hosted model fallbackReAct
4

Tooling / Action

Forced tool use to fetch ground-truth nutrition before committing a day.

function callingbind_toolstool_choice=submitlookup_usda_batch
Anti-hallucination guarantee

The LLM never writes nutrition numbers — the server computes them from USDA ground truth.

A hard architectural constraint: the submit schema has no nutrition fields, so the model cannot fabricate macros — stronger than any prompt instruction.

5

Knowledge / Retrieval (RAG)

Two complementary RAG systems — exact ingredient macros vs open-ended dish ideas.

Lexical USDA RAGfactual

Exact ingredient → macro grounding.

Postgres FTSts_rankLLM rerank
Semantic Recipe RAGcreative

Open-ended dish ideas before generation.

Qdrant~230k RecipeNLG recipestwo fused collections

two RAGs, two query shapes — exact ingredient macros vs open-ended dish ideas

Closed-loop data flywheel

Validated meals are re-ingested into the recipe corpus via a LlamaIndex pipeline — future generations retrieve from real past outputs, not just the seed corpus.

Drift guards prevent mode collapse: provenance weighting (×0.7), a ≤1-of-3 generated cap, semantic dedup (cosine >0.95), and a weekly variety audit with clean rollback.

6

Memory / State / Persistence

Typed graph state checkpointed to Postgres — an interrupted plan resumes from the next unfinished day.

MealPlanState (TypedDict + reducers)AsyncPostgresSaverTTL cache
7

Guardrails / Validation

Schema enforcement plus dietary-restriction checks on every output.

PydanticInstructorrestriction checker
8

Evaluation / Quality

A 3-layer online validate-and-retry cascade, backed by a pre-release eval harness.

programmatic checksLLM-as-judge (Haiku)USDA fact-checkoscillation detector
Pre-release eval harness — golden sets (101 meal plans + 80 chat cases) · RAGAS · DeepEval · MLflow · best score 0.971
SSE → validated meal-plan days
Cross-cutting concerns

9. Observability

cross-cutting · spans every layer

Traces every LLM call across the whole stack.

self-hosted Langfuse v3prompt cachingOpenTelemetry

10. Data / Infrastructure

cross-cutting · spans every layer

Runs on Hetzner + Cloudflare Workers (moved from Railway in August 2026).

HetznerCloudflare WorkersPostgreSQLRedisQdrant

AI System

The graph, tools and retrieval behind a meal plan. A LangGraph StateGraph generates the plan one day at a time, a ReAct tool loop grounds each day in USDA data, and two retrieval stores answer two different questions: which food is this, and what should we cook.

Meal-generation graph

LangGraph StateGraph · meal_generation.py
START
prepare_context

Builds the per-plan context in graph state before the first day.

retrieve_recipesconditional

Joins the graph when recipe RAG is on: pulls candidate dishes from Qdrant into state.

generate_day ⇄ usda_tools

One day of meals, produced by a tool-using ReAct loop.

generate_day
Claude reasons, picks tool calls
ReAct
usda_tools
executes lookups, returns observations

The loop ends when the model calls the forced submit tool DayPlanLLMSubmit (bind_tools(…, tool_choice=…)). Its schema has no nutrition fields.

deterministic compositor sums USDA values per ingredientportion solver scales grams· the LLM never writes a nutrition number
validate_day

3-layer check. On failure, a conditional edge sends structured feedback back to generate_day; the oscillation detector stops a day that flips between the same failures.

1 · programmatic checks2 · LLM judge3 · USDA fact-checkoscillation detector
pass
emit_day

Streams the finished day out as an SSE event.

update_history

Records the day, then a conditional edge loops to the next day or ends the run.

all days done
END
AsyncPostgresSaver checkpoints graph state after every node. Resume is day-granular: a restarted run starts at the next unfinished day.
Multi-participant plans: supervisor-workers graph
supervisorSend() fan-outworker × participantaggregatereconcile

Each participant is generated in parallel, then the results are merged and reconciled into one household plan (4+ participants).

Tools

Generation agentgenerate_day
  • lookup_usda_batch — Many ingredients in one call; schema-enforced batch input.
  • search_usda_food — Search USDA foods by name.
  • get_usda_nutrition — Nutrition values for one USDA food.
  • DayPlanLLMSubmit — Forced submit tool that ends the day. No nutrition fields.
Chat agent16 tools

A separate LangChain / LangGraph agent that answers questions and acts on the user's plans, e.g.:

list_participantsget_generated_plancreate_meal_planstart_meal_plan_generationlookup_nutrition_factssearch_recipescheck_plan_feasibility+ 9 more

RAG over two stores

USDA nutrition groundingfactual

Its own PostgreSQL 17. Maps an ingredient string to the right USDA food, so every macro traces back to USDA data.

  1. full-text search: to_tsvector · plainto_tsquery · ts_rank
  2. top 40 candidates
  3. Claude LLM reranker picks the match
Pick accuracy 0.64 → 0.92 on frozen probes
Recipe RAGcreative

Qdrant over ~230k filtered RecipeNLG recipes. Feeds dish ideas into retrieve_recipes; it shapes dish choice, never the numbers.

  1. tag pre-filter
  2. ANN search per collection · 2 collections
  3. RRF fusion across both result lists
  4. post-filter
  5. LLM rerank

State and serving

requestNext.jsNestJSBullMQ queueFastAPI AI service
SSEPythonNestJSNext.js· meals stream live, day by day
Anthropic ⇄ self-hosted model fallbackLangfuse v3 over OpenTelemetryAsyncPostgresSaver checkpointsHetzner + Cloudflare Workers

Key Features

  • pnpm monorepo with NestJS backend, Python FastAPI ai-service, and Next.js frontend
  • LangGraph StateGraph agent with server-computed USDA nutrition
  • Real-time SSE streaming from Python → NestJS → Next.js
  • Two retrieval jobs — USDA lookup (Postgres FTS + LLM rerank) and Qdrant recipe retrieval over ~230k RecipeNLG recipes
  • Closed-loop data flywheel re-ingests validated meals with drift guards (provenance weighting, ≤1/3 generated cap, semantic dedup, weekly variety audit)
  • Pre-release eval harness — golden sets, RAGAS/DeepEval/MLflow, best score 0.971
  • Crash-resume via Postgres-checkpointed graph state; self-hosted Langfuse observability
  • Live Anthropic ⇄ self-hosted model fallback; deployed on Hetzner + Cloudflare Workers

Tech Stack

Backend

NestJSFastAPIPythonPostgreSQLBullMQRedisPassport.js JWTSSE

AI Orchestration

LangGraphLangChainReActFunction CallingLangGraph Send()Claude Haiku 4.5 / SonnetSelf-Hosted LLM Fallback

RAG & Retrieval

QdrantRecipeNLG (~230k)RRF FusionPostgres FTSLLM Reranker

Eval & Observability

RAGASDeepEvalpromptfooMLflowGolden SetsLangfuse v3 (self-hosted)OpenTelemetryPrompt CachingInstructor

Frontend

Next.jsReactTypeScriptshadcn/uiSSR/SSG

Infrastructure

HetznerCloudflare WorkersDockerpnpm monorepoStripe

Challenges & Solutions

Preventing Hallucinated Nutrition Numbers

Problem

LLMs confidently fabricate plausible-but-wrong calorie and macro values, and a meal planner that lies about nutrition is worse than useless.

Solution

The LLM's submit schema (DayPlanLLMSubmit) has no nutrition fields, so the model physically can't write a macro number — it must call a USDA lookup tool for ground truth and the server computes every calorie and macro. A hard architectural constraint, not a prompt instruction.

Real-Time Streaming Across Services

Problem

Meal generation takes 10-30s through the AI pipeline. Users need immediate feedback, but the response crosses 3 service boundaries (Python → NestJS → Next.js).

Solution

SSE streaming pipeline where the FastAPI ai-service emits Server-Sent Events to NestJS, which proxies them to the Next.js frontend behind a configurable 360 s hang watchdog. Users see meals being generated in real-time.

Staying Up When the LLM Provider Fails

Problem

A production outage when the Anthropic account balance ran out showed that every generation depended on a single LLM provider.

Solution

Added a live cross-provider fallback — Anthropic ⇄ a self-hosted model — so generation switches providers instead of failing. It has fired in production when the Anthropic balance ran out.

Crash-Resumable Long-Running Generation

Problem

Generating a full multi-day plan is long-running, and a crash or timeout mid-plan would otherwise discard every day already produced.

Solution

Graph state is checkpointed to Postgres via AsyncPostgresSaver, keyed by thread_id, so an interrupted plan resumes by restarting from the next unfinished day instead of starting over.

Reliable AI Meal Generation

Problem

An LLM is non-deterministic, yet a multi-day meal plan needs a reliable, repeatable pipeline with tool use, retries, and streaming — a hand-rolled while-loop gets brittle fast.

Solution

Modeled generation as a LangGraph StateGraph: a deterministic day loop (prepare_context → generate_day ⇄ usda_tools → validate_day → emit_day → update_history) wraps a ReAct-style inner loop where the LLM decides tool calls, and a forced submit tool cleanly terminates each day.

Enforcing Quality on Non-Deterministic Output

Problem

A generated day can still miss calorie/macro targets or violate dietary restrictions (keto, vegan, allergen-free), and a single LLM pass can't be trusted.

Solution

A 3-layer validate-and-retry cascade — programmatic checks → optional LLM-as-judge (Haiku) → a blocking USDA ground-truth fact-check that retries at >50% calorie or >60% macro deviation — re-prompts with structured feedback injected into the next attempt, with an oscillation detector to bail on stuck days.

Two Retrieval Jobs: Nutrition vs Dish Ideas

Problem

Exact ingredient→macro lookup and open-ended dish inspiration are different retrieval problems. Short ingredient names ('chicken breast, raw') are won by lexical search, while conceptual dish queries ('Moroccan chickpea stew' ≈ 'North African lentil soup') are lexically distant but semantically equivalent, so full-text search misses them — one retrieval mechanism can't serve both.

Solution

Two complementary retrieval paths that don't compete. The factual layer is USDA lookup — Postgres full-text search plus an LLM rerank — grounding every macro. The creative layer is recipe retrieval on Qdrant over ~230k filtered RecipeNLG recipes, fusing two collections (~229k and ~2.1M points), whose results feed dish ideas into generation. USDA still grounds every number; recipe retrieval only shapes dish selection.

Preventing Mode Collapse in the Data Flywheel

Problem

Re-ingesting the model's own validated meals into the recipe corpus closes a data flywheel, but conditioning future generation on past generations risks an ever-narrowing dish distribution — mode collapse, where the system retrieves and re-generates the same handful of dishes in a tightening loop.

Solution

A guarded self-ingestion ETL: only meals that pass all 3 cascade layers with <5% deviation, no oscillation, and no force-commit qualify. Drift guards prevent collapse — provenance weighting (generated recipes ×0.7 at retrieval), a fraction cap (≤1 of 3 injected recipes may be generated), semantic dedup (skip if cosine >0.95), and a weekly variety audit with a provenance audit trail for clean rollback. Feature-flagged until 30-day variety metrics prove stable.

Key Achievements

pnpm Monorepo
3 services: NestJS, FastAPI, Next.js
LangGraph Agent
Tool-using agent with server-computed USDA nutrition
Crash-Resume
Postgres-checkpointed state for long-running generation
105 Users
~$0.26–0.33 per plan on Haiku 4.5 / Sonnet
Melio MealPlan AI - Project | Oleksandr Yusypenko