Backend & Data Architecture
Data platforms and the backends on top of them: Temporal-orchestrated ETL, CDC into a ClickHouse read model, and fast, cached APIs. 8+ years building scalable data systems, most recently over ~94M federal award transactions.
What I Build
From ETL and CDC to OLAP analytics and fast APIs — data architecture that scales and proves it is correct.
Data Pipelines & ETL
Build ingestion pipelines that prove they are complete. Most recently SmartSync at GovChime: SAM.gov ETL orchestrated by Temporal, with every dataset checked against a second copy of the source instead of trusting a green run.
- Temporal workflows: 32 schedules, 6 task queues
- Extract → transform → load → verify per window
- Completeness vs the source's own files (0 of 80,922 missing)
- API-key pools, rate budgets and usage ledgers
CDC, CQRS & Replication
Keep a write model and its read models in sync — and prove it. PostgreSQL as the system of record, a ClickHouse read model fed by watermark-based CDC, and logical-replication mirrors with reconciliation.
- Watermark-based CDC into ClickHouse
- Daily checksum parity (94.4M award transactions)
- PostgreSQL logical-replication mirrors
- Health and hash-reconciliation workflows
Microservice Design
Architect and build microservice systems with clear domain boundaries, event-driven communication, and independent deployability. From monolith decomposition to greenfield multi-service platforms.
- Domain-driven service boundaries
- Event-driven architecture
- API gateway patterns
- Independent deployment & scaling
OLAP & Analytics
Design ClickHouse OLAP solutions for real-time analytics on massive datasets. Materialized views, columnar storage optimization, and hybrid PostgreSQL + ClickHouse architectures for the best of both worlds.
- ClickHouse columnar analytics
- Materialized views + CQRS read projections
- ~94M award-transaction dataset handling
- Headline query cut from 13.6 s to ~1 s
Performance & Caching
Make data-heavy APIs and frontends fast: index-friendly queries, response caches with stampede protection, and Next.js on Cloudflare Workers with cache freshness matched to how often the data changes.
- Query rewrites measured with EXPLAIN ANALYZE
- Single-flight guards against cache stampedes
- Next.js on Cloudflare Workers (OpenNext)
- On-demand revalidation with tag caches
AI-Augmented Development
Leverage AI-native development workflows for rapid, high-quality delivery. Claude Code with custom agent skills, TDD-driven AI code generation, and MCP integrations for plan-implement-test-iterate loops.
- Claude Code & MCP integrations
- TDD-driven AI code quality
- Custom agent skills
- Plan-implement-test-iterate loops
Technologies I Use
Battle-tested backend stack for building scalable data systems
Backend Projects
Melio MealPlan AI
AI-Powered Meal Planning
NestJS + Python FastAPI dual-service backend with PostgreSQL, a BullMQ + Redis job queue, and Passport.js JWT auth. NestJS hands generation jobs to the FastAPI ai-service, which streams Server-Sent Events back across 3 service boundaries (Python → NestJS → Next.js) behind a configurable 360 s hang watchdog. Stripe handles subscription billing, and the platform runs on Hetzner + Cloudflare Workers.
GovChime Analytics Platform
Federal Contract Data Platform — ETL, CDC & CQRS
**Data platform first.** SmartSync, the SAM.gov + USAspending.gov ETL, runs on Temporal in production since 2026-09-25 — 32 schedules on 6 task queues, extract → transform → load → verify per window, a 13-key SAM.gov pool with a usage ledger. Completeness is checked against SAM.gov's own files: 0 of 80,922 active notices missing. **CDC and CQRS.** PostgreSQL is the write model; a Temporal coordinator per table projects new rows into the ClickHouse read model by watermark, with daily checksum parity over ~94.4M award transactions. A logical-replication PostgreSQL mirror runs as Temporal Workflows with hourly health and daily hash reconciliation. **Serving it fast.** Express 4 + Sequelize (TypeScript) API with ~242 endpoints and 78 filters: a headline query 13.6 s → ~1 s, ClickHouse queries 2.2 s → 0.27 s, a single-flight stampede guard on the in-memory response cache, cache TTLs aligned at 10 min, and a ClickHouse → Postgres fallback chain later simplified (32 Postgres MVs dropped, ClickHouse MVs 31 → 14). Running on OVHcloud with Docker and Komodo.
Filament Web3 Airdrop Platform
Token Distribution & Delegate Voting on an L2 Rollup
Node server routes that aggregate an external blockchain-analytics API for the campaign dashboard with per-widget graceful degradation, and the integration with the team's L2 rollup REST API — queries, sequencer submission and ledger status. Rollup transactions are signed in MetaMask through the rollup's own serializer compiled to WebAssembly, with BigInt-safe payloads and a nonce-0 fix for wallets the rollup has not seen yet.
Primsell NFT E-Commerce
Node.js Business Logic on PostgreSQL — Checkout, Stripe Connect & Royalty Ledger
Backend business logic in Node.js on PostgreSQL: an Express + Sequelize core API (~79 REST endpoints) and a NestJS + MikroORM payment service on Stripe Connect. Checkout overselling fixed with SELECT … FOR UPDATE and atomic SQL counters; Stripe late-payment and float-money edge cases handled; a royalty ledger fed by an OpenSea/Rarible resale indexer with confirmation-gated withdrawals; idempotent burn-to-redeem with a recovery job; a PostgreSQL trigger plus backfill for CRM data.
Employee Engagement Platform
Multi-Tenant Gamification Backend on Python FastAPI
Python FastAPI backend for a multi-tenant gamification platform for hourly workers: subdomain tenancy, JWT access/refresh with role guards, an append-only points ledger with approvals, gift-card redemption, KPI goal and insights endpoints, a WebSocket notifications inbox, and a UKG Pro integration driving shift polls and SMS. ~73 API operations used by the web client.
Hospital Benchmarking Report Engine
Express Report Service & Server-Driven React Report UI
Express report-engine service for a hospital benchmarking platform: an entitlement-scoped token exchange with the main platform, report and dashboard endpoints with drill-downs, and the KPI business logic — each hospital's indicators against its peer group's median and quartiles, with multi-year trends. Responses are typed UI blocks, so new reports ship without a frontend release.
SpaceSeven NFT Marketplace
EVM + Concordium Indexer Drivers & a Shared TypeScript Marketplace Layer
Chain-indexer network drivers for Ethereum and Concordium: per-contract block cursors seeded at each deployment block (backfill on cold start, resume on restart), contract events decoded into one chain-neutral protobuf event with 23 types, and HMAC-SHA256-signed delivery to the marketplace backend, which applies each event idempotently in the same transaction as the state change.
ROBBED_
Memecoin Launchpad on an Arbitrum Orbit L2
**Indexer.** Ponder over the on-chain event families → Postgres with `pg_trgm`, the single source of derived truth: venue-continuous candles across six intervals, `Transfer`-sourced holder balances, confirmation-state watermarks, metadata-hash verification, and creator-fee accrual — one Redis publish per handler and zero hot-path reads. **API + WS.** Hono on Bun as two processes — HTTP (25+ read endpoints over indexer tables, `pg_trgm` search, API-mediated R2 uploads, server-side metadata canonicalization, moderation gating, SIWE admin, per-token OG rendering via satori + resvg) and a Bun WebSocket fanout relaying Redis to sockets. The API never writes to the chain. **Keeper.** A small Bun + viem service that makes graduation automatic — a topic-filtered `eth_subscribe` on `GraduationReady` fires the permissionless `graduate()` within ~1–2 blocks, with a Postgres sweep as the fallback for WS drops, an on-chain `phase()` re-read before every send for idempotency, and a cooldown that stops persistent-revert hot-loops. It holds no privileged role and adds zero new authority.
Backend Architecture FAQ
When should I use ClickHouse vs PostgreSQL?
PostgreSQL is excellent for transactional workloads (CRUD, user data, business logic). ClickHouse shines for analytical queries on large datasets — aggregations, time-series, reporting dashboards. I often use both: PostgreSQL as the source of truth and ClickHouse for OLAP analytics, with materialized views bridging the two.
How do you make sure a data pipeline is complete?
A green run only proves the code executed. I check each dataset against a second copy of the source. At GovChime the Temporal-orchestrated SAM.gov pipelines are compared daily with SAM.gov's own notice files (0 of 80,922 active notices missing), and ClickHouse is compared with PostgreSQL by checksum every day. A gap triggers a repair run instead of a silent pass.
How do you approach microservice architecture?
I start with domain-driven design to identify service boundaries. Each service owns its data and communicates via events or APIs. I use NestJS for structured backend services, Redis for caching, and RabbitMQ for async messaging. The goal is independent deployability without premature complexity.
What does AI-augmented development mean in practice?
I use Claude Code as my primary development partner with custom MCP integrations and agent skills. The workflow is: write test specs first (TDD), then use AI to generate implementation, validate against tests, and iterate. This delivers 2-3x development speed while maintaining code quality through automated testing.
Can you work with my existing backend?
Yes. I regularly integrate into existing Node.js/NestJS codebases. Whether it's adding ClickHouse for analytics, decomposing a monolith into services, or optimizing slow queries — I can work incrementally without disrupting your current system.
Hiring a Backend / Data Engineer?
I'm open to B2B contracts and full-time senior engineering roles. Reach out and let's talk about your data architecture challenges.
Get in Touch