Running NestJS on EC2 with systemd
By Isaiah Daniel·
A systemd unit for a NestJS API on EC2: env file secrets, restart policy, graceful shutdown, journald logs and symlinked releases for near-zero-downtime deploys.
On one backend I inherited, the production API ran inside a screen session that someone had started months earlier. When the instance rebooted for a kernel patch, the API simply did not come back. Nobody noticed until a partner's webhook retries started piling up. The fix was not Kubernetes or a new platform. It was about forty lines of systemd configuration and a deploy script that stopped treating the server like a laptop.
Not every service needs a container orchestrator. A single NestJS API on one or two EC2 instances, behind a load balancer or nginx, is a perfectly reasonable setup, as long as the process is supervised properly. This post is the setup I reach for, using Brightpath HR, a fictional payroll SaaS, as the example.
What systemd gives you for free
systemd is already running on every mainstream Linux AMI (Amazon Linux 2023, Ubuntu). Used properly, it replaces pm2, forever and hand-rolled nohup scripts:
- Starts the API on boot, in the right order relative to networking.
- Restarts it when it crashes, with rate limiting so a crash loop does not spin forever.
- Sends a clean
SIGTERMon stop and escalates toSIGKILLafter a timeout you choose. - Captures stdout and stderr in journald, with no log files for you to rotate.
- Runs the process as an unprivileged user with filesystem sandboxing.
The directory layout
Before the unit file, decide where things live. I use a releases layout borrowed from Capistrano-style deploys:
/srv/brightpath-api/
releases/
20261004T101500/
20261005T093000/
current -> releases/20261005T093000
shared/ # uploads, anything that must survive a deploy
/etc/brightpath-api/env # secrets, root-owned, 0600
The service always runs from current. A deploy builds a new release directory, flips the symlink, and restarts. Rollback is flipping the symlink back.
The unit file
/etc/systemd/system/brightpath-api.service:
[Unit]
Description=Brightpath HR API
After=network-online.target
Wants=network-online.target
StartLimitIntervalSec=60
StartLimitBurst=5
[Service]
Type=simple
User=brightpath
Group=brightpath
WorkingDirectory=/srv/brightpath-api/current
EnvironmentFile=/etc/brightpath-api/env
ExecStart=/usr/bin/node dist/main.js
Restart=on-failure
RestartSec=5
KillSignal=SIGTERM
TimeoutStopSec=30
LimitNOFILE=65536
SyslogIdentifier=brightpath-api
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/srv/brightpath-api/shared
[Install]
WantedBy=multi-user.target
Then:
sudo systemctl daemon-reload
sudo systemctl enable --now brightpath-api
systemctl status brightpath-api
A few lines deserve explanation, because each one is there because of a failure I have seen.
ExecStart runs node directly, not npm run start:prod. npm adds a parent process, extra memory, and noisy exit codes. With node as the main process, systemd's view of "is the service alive" matches reality. Use an absolute path: if node was installed with nvm under someone's home directory, /usr/bin/node will not exist and ProtectHome=true will hide it anyway. Install node system-wide (NodeSource packages or the distro package).
Restart=on-failure with start limits. On a non-zero exit or a crash signal, systemd waits RestartSec and starts it again. StartLimitBurst=5 within StartLimitIntervalSec=60 means five failed starts in a minute puts the unit into a failed state instead of looping. That is what you want when a bad environment variable makes the app die on boot: it fails loudly, and your alerting catches a failed unit rather than a flapping one. Note that these two settings live in [Unit] on modern systemd; in [Service] they are ignored with a warning.
TimeoutStopSec=30. On stop or restart, systemd sends SIGTERM, waits up to 30 seconds, then sends SIGKILL. Your app's graceful shutdown must finish inside that window, so size it to your slowest legitimate request or job.
The sandboxing block. ProtectSystem=strict makes the whole filesystem read-only for the process except paths you list in ReadWritePaths. If your app writes temp files or uploads somewhere unexpected, you find out on day one rather than after a compromise.
EnvironmentFile gotchas
/etc/brightpath-api/env looks like a .env file but is not parsed by a shell:
NODE_ENV=production
PORT=3000
DATABASE_URL=postgres://brightpath:secret@db.internal:5432/brightpath
REDIS_URL=redis://cache.internal:6379
- No
exportkeyword, and no$OTHER_VARexpansion. Write the final values. - Keep it
root:rootwith mode0600. systemd reads it as root before dropping to thebrightpathuser, so the app user never needs read access to the file itself. - Changing the env file needs
systemctl restart. Changing the unit file needsdaemon-reloadfirst. Mixing these up is the classic "I changed it and nothing happened."
If you would rather not keep secrets on disk at all, render this file at boot or deploy time from SSM Parameter Store using the instance role. The unit file does not change either way.
Graceful shutdown in NestJS
systemd sends SIGTERM, but NestJS ignores it unless you opt in:
// main.ts
async function bootstrap() {
const app = await NestFactory.create(AppModule);
app.enableShutdownHooks();
await app.listen(Number(process.env.PORT ?? 3000), '127.0.0.1');
}
bootstrap();
With shutdown hooks enabled, SIGTERM triggers app.close(): the HTTP server stops accepting new connections, in-flight requests finish, and lifecycle hooks run on your providers. That is where you close anything holding work:
@Injectable()
export class PayoutWorker implements OnApplicationShutdown {
private readonly worker = new Worker(
'payouts',
(job) => this.process(job),
{ connection: { url: process.env.REDIS_URL } },
);
async onApplicationShutdown(signal?: string) {
// Waits for the active job to finish before resolving.
await this.worker.close();
}
private async process(job: Job) {
// ...
}
}
The failure mode to watch is a payout job that takes longer than TimeoutStopSec. systemd will kill it mid-flight. That is survivable only if the job is idempotent, which is the same argument I made in the idempotent payments post: assume any process can die at any line, and make retries safe.
Two smaller gotchas: HTTP keep-alive connections can hold server.close() open until they idle out, so keep the stop timeout as a hard backstop. And binding to 127.0.0.1 means only nginx on the same host can reach the app, which is usually what you want.
Logs with journald
Because the app writes to stdout, logs are already captured:
journalctl -u brightpath-api -f # follow
journalctl -u brightpath-api --since "15 min ago"
journalctl -u brightpath-api -p err -b # errors since boot
Log JSON lines (pino or Nest's logger with a JSON formatter) so they stay greppable now and parseable when you ship them to CloudWatch or Loki later. On some AMIs journald keeps logs in memory only, so set Storage=persistent and a cap like SystemMaxUse=1G in /etc/systemd/journald.conf if you want them to survive a reboot.
Deploys with symlinked releases
The deploy script runs on the instance (from CI over SSH or SSM Run Command) with a prebuilt artifact:
#!/usr/bin/env bash
set -euo pipefail
APP=/srv/brightpath-api
REL="$APP/releases/$(date +%Y%m%dT%H%M%S)"
PREV=$(readlink -f "$APP/current" || true)
mkdir -p "$REL"
tar -xzf /tmp/brightpath-api.tgz -C "$REL"
(cd "$REL" && npm ci --omit=dev)
ln -sfn "$REL" "$APP/current.next"
mv -T "$APP/current.next" "$APP/current"
sudo systemctl restart brightpath-api
for i in $(seq 1 20); do
if curl -fsS http://127.0.0.1:3000/health > /dev/null; then
ls -1dt "$APP"/releases/* | tail -n +6 | xargs -r rm -rf
echo "Deployed $REL"
exit 0
fi
sleep 2
done
echo "Health check failed, rolling back to $PREV"
ln -sfn "$PREV" "$APP/current.next"
mv -T "$APP/current.next" "$APP/current"
sudo systemctl restart brightpath-api
exit 1
Why each piece matters:
mv -Tfor the swap.ln -sfndirectly oncurrentremoves and recreates the link, leaving a tiny window where it does not exist. Renaming a new link over the old one is atomic.- systemd resolves
WorkingDirectoryat start. The running process keeps its old directory until restart, which is why the restart comes after the swap, and why you never delete the release that is currently running. The script keeps the five newest. - Health check, then rollback. A deploy is not done when the process starts; it is done when
/healthanswers. Make that endpoint check the database and Redis connections, not just return 200. - Least-privilege sudo. The deploy user gets a sudoers line for
systemctl restart brightpath-apiand nothing else.
Getting closer to zero downtime
A single instance restart leaves a gap of a few seconds. If that matters, use a template unit, brightpath-api@.service, with Environment=PORT=%i, and run brightpath-api@3001 and brightpath-api@3002 behind nginx with both as upstreams. Restart one, wait for its health check, then the other. nginx's proxy_next_upstream retries idempotent requests against the healthy instance while one is down. Across multiple instances, a load balancer with connection draining does the same job one level up.
Takeaways
| Concern | Setting or practice |
|---|---|
| Survive reboots | enable the unit, WantedBy=multi-user.target |
| Crash recovery | Restart=on-failure plus StartLimitBurst in [Unit] |
| Secrets | Root-owned EnvironmentFile, 0600, no shell syntax |
| Clean stops | enableShutdownHooks(), close workers, size TimeoutStopSec |
| Logs | stdout as JSON, journalctl -u, persistent journald |
| Deploys | Release directories, atomic symlink swap, health check, rollback |
None of this is novel, which is the point. A boring, supervised process with a predictable deploy beats a clever setup that nobody remembers how to restart at 2 a.m.