Why I Moved AI Out of NestJS and Into a Dedicated Python LangGraph Service
When your AI pipeline lives inside your main backend, everything feels fine — until it doesn't. I hit that wall on MealPlan AI (now Melio), an AI-powered meal planning platform I work on as the main developer in a small team. NestJS handled both business logic and LLM orchestration.
The AI SDK chain was a black box. No observability into token costs. No crash recovery mid-generation. No way to validate LLM outputs before persisting them. When a 7-day meal plan generation failed on day 5, the entire thing restarted from scratch.
Something had to change.
The Architecture Split
I separated the system into two services with clear responsibilities:
- NestJS — business logic, auth, payments, job orchestration via BullMQ
- Python FastAPI — all LLM orchestration, validation, and AI-specific tooling
Why Python? LangGraph, LangChain, and the broader AI ecosystem are Python-first. Fighting that with TypeScript wrappers added complexity without adding value.
The 5-Node LangGraph StateGraph
The core of the Python service is a LangGraph StateGraph with five nodes:
- prepare_context — loads dietary restrictions, participant profiles, calorie targets, and retrieves relevant recipes via RAG (2.2M recipes from RecipeNLG, 80K foods from USDA)
- generate_day — calls the LLM with structured prompts including diversity history and calorie distribution targets
- validate_day — two-layer validation: programmatic restriction checking first, optional LLM self-validation second
- emit_day — streams a DAY_COMPLETED SSE event so NestJS can persist immediately via JSONB atomic append
- update_history — tracks dishes, ingredients, and cuisines to enforce diversity across the full meal plan
AsyncPostgresSaver checkpointing lets us resume from exactly where we stopped — no wasted LLM calls.
Type Safety Across Languages
A dual-language service creates a type drift risk. I solved this with a one-directional pipeline: Zod schemas (TypeScript) export to JSON Schema, which generates Pydantic models (Python). A CI workflow runs on every PR to catch drift before it reaches production.
What This Unlocked
Splitting AI into its own service wasn't just a refactor — it enabled features that would have been painful to build in the monolithic setup:
- Langfuse observability — every LLM call traced with token counts, costs, and latency. Self-hosted, full control over data.
- Incremental persistence — each day saves immediately. A
PARTIALLY_COMPLETEDstatus lets users see progress and resume interrupted plans. - Granular regeneration — separate endpoints for regenerating a single day (full graph) or a single meal (direct LLM call) with user feedback injected into prompts.
- RAG retrieval — full-text search against RecipeNLG and USDA datasets with pgvector fallback for semantic search.
- 329 Python tests — pytest-asyncio covering every node, validator, and edge case independently from the NestJS test suite.
Key Takeaways
- Separate AI from business logic early. The longer you wait, the harder the extraction. AI services have different scaling, testing, and deployment needs.
- Use LangGraph for multi-step AI workflows. A linear chain breaks down when you need validation loops, conditional retries, and state management across steps.
- Invest in type contracts across languages. Zod-to-JSON-Schema-to-Pydantic catches bugs at build time that would otherwise surface as silent data corruption.
- Stream incrementally, persist incrementally. Users shouldn't wait for a 30-second generation to complete before seeing anything. SSE + atomic JSONB appends make this straightforward.

Oleksandr Yusypenko
Senior Full-Stack + AI Engineer. Building in public — AI agents, LangGraph, production systems.
Related Posts
Jul 13, 2026
Claude Code config is a team asset, not a personal file
On the robbed monorepo I set up Claude Code the way you'd set up CI or CODEOWNERS: a shared, version-controlled, enforced config so every engineer and every AI agent inherits the same conventions, ownership boundaries, and guardrails.
Jul 8, 2026
We Gave Our Agent a Third Tool. Then We Deleted It.
Melio's meal-plan agent runs on exactly two tools. We eval-tested a third — a deterministic self-check — and it tripled cost per plan and made accuracy worse. Here's the data that got it deleted.
Jul 8, 2026
The Prompt Cache That Silently Did Nothing
We shipped Anthropic prompt caching with green unit tests — and it did nothing in production. The culprit: model-specific cache floors and per-day content ahead of the breakpoint. Here's the diagnosis and the fix.
Agent Builders Are Changing How I Ship Code — Here's My Actual Workflow
NextPostgreSQL + ClickHouse: The Dual-Database Pattern That Made 90M-Row Dashboards Instant