ID

Home·Writings

Node.js + Python Architecture for AI Applications

By ·

Splitting an AI product between a Node.js API and a Python service: who owns what, HTTP vs queues, timeouts, shared contracts, and when not to split.

On a health and nutrition product I worked on, the first version of the recommendation feature shipped as a FastAPI service that the mobile app called directly. It worked in the demo. Within a month we had two auth implementations drifting apart, a Python refactor that renamed protein to protein_g and quietly blanked out a screen in the app, and no single place to see why a given user got a given meal plan.

The fix was not a different framework. It was deciding, explicitly, which runtime owns which job, and putting a contract between them. This post is the architecture I now use when a product needs both a Node.js backend and Python for the model work.

Who owns what

The rule I settled on is simple: Node owns the product, Python owns the math.

ConcernNode.js (NestJS)Python (FastAPI)
Auth, sessions, usersYesNo
Public API and rate limitsYesNo
Product data (Postgres)Owns writesReads only what it is sent
Models, features, scoringCalls itYes
Caching and fallbacksYesNo

The Python service is never exposed to the internet. The mobile and web clients talk to Node, and Node talks to Python over a private network. That one decision removed the duplicate auth, gave us one place to log every recommendation, and meant the Python side could change shape without the app noticing, as long as the contract held.

The Python side: small, typed, stateless

In that product, the core calculation was the classic one: basal metabolic rate (Mifflin-St Jeor), total daily energy expenditure from an activity factor, then calorie and macro targets for the user's goal. Ranking meals came later, but it sat behind the same kind of endpoint.

import hmac
import os
from typing import Literal

from fastapi import APIRouter, Depends, FastAPI, Header, HTTPException
from pydantic import BaseModel, Field

TOKEN = os.environ["RECS_TOKEN"]
MODEL_VERSION = "targets-2026.10"

ACTIVITY = {"sedentary": 1.2, "light": 1.375, "moderate": 1.55, "active": 1.725, "very_active": 1.9}
GOAL_ADJUST = {"lose": -500, "maintain": 0, "gain": 300}


def internal_only(authorization: str = Header(default="")) -> None:
    if not hmac.compare_digest(authorization, f"Bearer {TOKEN}"):
        raise HTTPException(status_code=401)


class ProfileIn(BaseModel):
    sex: Literal["male", "female"]
    age: int = Field(ge=13, le=100)
    weight_kg: float = Field(gt=20, lt=400)
    height_cm: float = Field(gt=100, lt=250)
    activity: Literal["sedentary", "light", "moderate", "active", "very_active"]
    goal: Literal["lose", "maintain", "gain"]


class Targets(BaseModel):
    bmr: int
    tdee: int
    calories: int
    protein_g: int
    carbs_g: int
    fat_g: int
    model_version: str


router = APIRouter(prefix="/v1", dependencies=[Depends(internal_only)])


@router.post("/targets", response_model=Targets)
def targets(p: ProfileIn) -> Targets:
    s = 5 if p.sex == "male" else -161
    bmr = 10 * p.weight_kg + 6.25 * p.height_cm - 5 * p.age + s
    tdee = bmr * ACTIVITY[p.activity]
    calories = tdee + GOAL_ADJUST[p.goal]
    protein = 1.8 * p.weight_kg
    fat = 0.25 * calories / 9
    carbs = max((calories - protein * 4 - fat * 9) / 4, 0)
    return Targets(
        bmr=round(bmr), tdee=round(tdee), calories=round(calories),
        protein_g=round(protein), carbs_g=round(carbs), fat_g=round(fat),
        model_version=MODEL_VERSION,
    )


app = FastAPI()
app.include_router(router)


@app.get("/healthz")
def healthz() -> dict:
    return {"ok": True}

A few things here are deliberate:

  • def, not async def. FastAPI runs plain def handlers in a thread pool. CPU-bound work inside an async def blocks the event loop and stalls every other request on that worker. Reserve async def for handlers that actually await I/O.
  • Auth on the router, not the app. The load balancer health check hits /healthz without a token. Putting the dependency on the whole app makes every health check a 401 and the target gets marked unhealthy.
  • model_version in every response. When someone asks "why did this user get 1,400 calories last Tuesday?", the answer starts with which version of the logic produced it. Node stores it next to the result.
  • Stateless. Python gets everything it needs in the request. It does not reach into the product database, which keeps the ownership line clean and makes the service trivial to scale horizontally.

The Node side: timeouts, contracts, fallbacks

Pydantic guarantees what Python sends. Nothing guarantees what Node receives, unless Node checks. I validate the response with zod at the boundary, so contract drift becomes a loud, logged error instead of a blank screen.

import { Injectable, Logger, ServiceUnavailableException } from '@nestjs/common';
import { z } from 'zod';

const Targets = z.object({
  bmr: z.number().int(),
  tdee: z.number().int(),
  calories: z.number().int(),
  protein_g: z.number().int(),
  carbs_g: z.number().int(),
  fat_g: z.number().int(),
  model_version: z.string(),
});
export type Targets = z.infer<typeof Targets>;

export type ProfileInput = {
  sex: 'male' | 'female';
  age: number;
  weight_kg: number;
  height_cm: number;
  activity: 'sedentary' | 'light' | 'moderate' | 'active' | 'very_active';
  goal: 'lose' | 'maintain' | 'gain';
};

@Injectable()
export class RecommendationClient {
  private readonly logger = new Logger(RecommendationClient.name);
  private readonly baseUrl = process.env.RECS_URL!;
  private readonly token = process.env.RECS_TOKEN!;

  async targets(profile: ProfileInput, requestId: string): Promise<Targets> {
    let res: Response;
    try {
      res = await fetch(`${this.baseUrl}/v1/targets`, {
        method: 'POST',
        headers: {
          'content-type': 'application/json',
          authorization: `Bearer ${this.token}`,
          'x-request-id': requestId,
        },
        body: JSON.stringify(profile),
        signal: AbortSignal.timeout(800),
      });
    } catch (err) {
      this.logger.warn(`recs unreachable (${(err as Error).name}) req=${requestId}`);
      throw new ServiceUnavailableException('Recommendations unavailable');
    }

    if (!res.ok) {
      this.logger.error(`recs returned ${res.status} req=${requestId}`);
      throw new ServiceUnavailableException('Recommendations unavailable');
    }

    const parsed = Targets.safeParse(await res.json());
    if (!parsed.success) {
      this.logger.error(`recs contract drift req=${requestId}: ${parsed.error.message}`);
      throw new ServiceUnavailableException('Recommendations unavailable');
    }
    return parsed.data;
  }
}

The caller decides what a failure means for the user. For targets, a slightly old answer is far better than an error, so the service falls back to the last stored result and says so:

async getPlan(userId: string, requestId: string) {
  const profile = await this.profiles.get(userId);
  try {
    const targets = await this.recs.targets(profile, requestId);
    await this.plans.saveLatest(userId, targets);
    return { ...targets, stale: false };
  } catch (err) {
    const cached = await this.plans.latest(userId);
    if (cached) return { ...cached, stale: true };
    throw err;
  }
}

Gotchas worth calling out:

  • Pick the timeout from the user's budget, not the service's average. If the screen should load in under a second, Python gets 800ms, not "whatever it takes". AbortSignal.timeout makes this a one-liner in Node 18+.
  • Retry reads, be careful with writes. A pure calculation is safe to retry once with jitter. If the Python endpoint writes anything (a stored embedding, a training example), give it an idempotency key first. I covered that pattern in idempotent payments with Redis and PostgreSQL, and it applies unchanged here.
  • A 422 from Python is a Node bug. Node already validated the profile. If FastAPI rejects it, the two schemas disagree, and that should page someone rather than show the user a validation message.
  • Propagate the request ID. One x-request-id across both services turns "the plan was wrong" into a single log search.

HTTP or a queue?

Synchronous HTTP is right when a user is waiting and the work finishes in well under a second: scoring, ranking a short list, computing targets. It is wrong when:

  • the work takes seconds or minutes (generating a full 30-day plan, batch re-scoring every user after a model change);
  • spikes would overwhelm the Python workers and you would rather queue than shed load;
  • the result can arrive later through a push notification or a status endpoint.

For those, Node writes a job record, enqueues a message, and returns 202 Accepted with the job ID. A Python worker consumes it, does the work, and calls back to Node (or writes to a results queue). For cross-language queues I prefer SQS or another broker with first-class clients in both languages over a Node-specific job library, so neither side has to reverse-engineer the other's job format.

The rule of thumb: if you are tempted to raise the HTTP timeout past a couple of seconds, you want a queue.

When not to split

Two services means two deploys, two sets of dependencies, a network hop and a contract to maintain. That cost is worth it only when you actually need Python:

  • Calling a hosted LLM API? Stay in Node. The official SDKs are just as good, and a second service adds latency for nothing.
  • Arithmetic like BMR and TDEE alone? Honestly, that could live in Node too. The split paid off for us once ranking needed pandas and scikit-learn, and once the person iterating on the model worked in notebooks.
  • numpy, pandas, scikit-learn, PyTorch, or a data scientist shipping the model? Split, and use the boundary above.

Takeaways

  • Node owns users, auth, product data and the public API. Python owns models and is never public.
  • Validate on both sides of the wire: pydantic on the way out, zod on the way in.
  • Return a model_version with every result and store it.
  • Set timeouts from the user's latency budget, and decide the fallback before you need it.
  • Use HTTP when a user is waiting and the work is fast; use a queue for anything slow, bursty or batch.
  • Do not split just because the feature is called "AI". Split when the Python ecosystem is doing real work.