Building an RTMP-to-HLS Video Pipeline on AWS
Taking a live RTMP feed from an encoder to adaptive HLS on CloudFront: MediaLive vs nginx-rtmp on EC2, segment sizing, caching, and the latency vs cost trade.
On a media platform I worked on, video was strictly on demand: an editor uploaded a file, MediaConvert produced an HLS ladder, and CloudFront served it. That pipeline was calm because every step had a finished file to work with. Then the inevitable request came in: "Can we go live for the election results?"
Live is a different animal. There is no finished file. The encoder pushes a continuous RTMP stream, every second of delay is visible to viewers, and if the ingest box falls over, the broadcast is gone rather than just late. This post walks through how I would build the live leg for Northlight Press, a fictional newsroom, and how it plugs into the on-demand pipeline that already exists.
The shape of the problem
The newsroom's encoder (OBS, vMix, or a hardware box in the studio) speaks RTMP. Browsers and phones do not play RTMP anymore; they play HLS: a playlist (.m3u8) that points at short media segments (.ts or fMP4), refreshed every few seconds. So the pipeline has four jobs:
- Ingest an authenticated RTMP push.
- Transcode it into an adaptive ladder (for example 720p and 480p).
- Segment each rendition into HLS and keep a rolling live playlist.
- Deliver it through a CDN without the origin melting when the audience arrives.
On AWS there are two realistic ways to do steps 1 to 3.
Option A: MediaLive (managed)
AWS Elemental MediaLive accepts an RTMP push input, transcodes to whatever ladder you define, and writes HLS to S3 or to MediaPackage for packaging and origin duties. You get things that are hard to build yourself:
- Redundancy. A
STANDARDchannel runs two pipelines in separate Availability Zones.SINGLE_PIPELINEis cheaper and fine for rehearsals. - Input security groups that only accept pushes from your studio's IP ranges.
- Archive outputs that drop the full broadcast into S3 for replay.
The catch is the billing model: you pay per hour that a channel is running, whether ten people or ten thousand are watching, and an idle channel left running over a weekend is a real line on the bill. If you choose MediaLive, make starting and stopping channels part of the event workflow (an API call from the admin panel when the producer clicks "Go live"), not something someone remembers to do in the console.
Option B: nginx-rtmp and ffmpeg on EC2 (self-hosted)
For a newsroom that goes live a few times a month, a single compute-optimized EC2 instance running nginx with the RTMP module is often enough. nginx accepts the push, hands it to ffmpeg for the ladder, and writes HLS to local disk:
rtmp {
server {
listen 1935;
chunk_size 4096;
application ingest {
live on;
record off;
on_publish http://127.0.0.1:3000/live/auth;
exec ffmpeg -i rtmp://127.0.0.1/ingest/$name
-c:v libx264 -preset veryfast -r 30 -g 60 -keyint_min 60 -sc_threshold 0
-b:v 2800k -s 1280x720 -c:a aac -b:a 128k -f flv rtmp://127.0.0.1/hls/$name_720p
-c:v libx264 -preset veryfast -r 30 -g 60 -keyint_min 60 -sc_threshold 0
-b:v 1200k -s 854x480 -c:a aac -b:a 96k -f flv rtmp://127.0.0.1/hls/$name_480p;
}
application hls {
live on;
allow publish 127.0.0.1;
deny publish all;
hls on;
hls_path /var/www/hls;
hls_nested on;
hls_fragment 2s;
hls_playlist_length 12s;
hls_variant _720p BANDWIDTH=2928000,RESOLUTION=1280x720;
hls_variant _480p BANDWIDTH=1296000,RESOLUTION=854x480;
}
}
}
A few details in there matter more than they look:
-g 60 -keyint_min 60 -sc_threshold 0at 30 fps forces a keyframe every 2 seconds, exactly matchinghls_fragment 2s. Segments can only be cut on keyframes, so if the GOP and the fragment length disagree, you get uneven segments and players that stall or drift between renditions.- The
hlsapplication only accepts publishes from localhost. Only ffmpeg should be able to write renditions; the outside world talks toingest. hls_variantmakes nginx write a master playlist at/hls/<key>.m3u8that lists both renditions, which is what the player loads.
Authenticating the push
on_publish makes nginx call your API before accepting a stream. It sends a form-encoded body including app and name (the stream key); any 2xx allows the publish and anything else rejects it. In NestJS that is a few lines:
import { Body, Controller, ForbiddenException, HttpCode, Post } from '@nestjs/common';
import { StreamsService } from './streams.service';
@Controller('live')
export class LiveAuthController {
constructor(private readonly streams: StreamsService) {}
@Post('auth')
@HttpCode(204)
async authorize(@Body() body: { app?: string; name?: string }): Promise<void> {
if (body.app !== 'ingest' || !body.name) throw new ForbiddenException();
const stream = await this.streams.findScheduledByKey(body.name);
if (!stream) throw new ForbiddenException();
await this.streams.markLive(stream.id);
}
}
Issue a fresh stream key per event and expire it afterwards. Keys get pasted into group chats and screenshots; a key that only works during a scheduled window limits the damage. Pair it with on_publish_done to mark the stream ended.
Delivery: CloudFront in front of the origin
Serving viewers straight from the EC2 box does not survive a traffic spike. Put CloudFront in front with the instance as a custom origin, and make caching behave differently for playlists and segments:
location /hls/ {
root /var/www;
types {
application/vnd.apple.mpegurl m3u8;
video/mp2t ts;
}
location ~ \.m3u8$ { add_header Cache-Control "max-age=1"; }
location ~ \.ts$ { add_header Cache-Control "max-age=86400"; }
}
| Object | Changes? | Cache |
|---|---|---|
.m3u8 playlists | Every segment (2s) | 1 second, so viewers see the live edge |
.ts segments | Never, once written | Long, so the origin serves each segment roughly once per edge |
Configure the CloudFront cache policy to honor origin headers (or set matching TTLs per path pattern). The classic bug is a default TTL that caches the playlist for minutes, so viewers sit on a frozen live edge while the encoder is fine.
Lock the origin so only CloudFront can reach it. AWS publishes a managed prefix list for exactly this:
data "aws_ec2_managed_prefix_list" "cloudfront" {
name = "com.amazonaws.global.cloudfront.origin-facing"
}
resource "aws_security_group_rule" "hls_from_cloudfront" {
type = "ingress"
from_port = 80
to_port = 80
protocol = "tcp"
prefix_list_ids = [data.aws_ec2_managed_prefix_list.cloudfront.id]
security_group_id = aws_security_group.live_origin.id
}
That prefix list counts as many entries against the security group's rule quota, so give it its own group. Restrict port 1935 to the studio's IPs separately where you can.
Latency vs cost
Glass-to-glass latency for HLS is roughly: encoder buffer, plus segment duration times the number of segments the player buffers before starting (commonly about three), plus CDN and network time. With 6-second segments you land around 20 to 30 seconds behind real time. With 2-second segments it drops to under ten, at the price of more requests per viewer and more playlist refreshes hitting CloudFront.
For a results night, ten seconds is fine. If you genuinely need two or three seconds (live betting, auctions, two-way interviews), that is Low-Latency HLS or WebRTC, and a different design conversation.
On cost, the two options fail differently. MediaLive costs the same per hour regardless of audience but buys redundancy. A single EC2 box is cheap but is a single point of failure: if it dies mid-broadcast, the stream is over until it comes back. In both cases, once the audience grows, CloudFront egress dominates the bill, which is another reason to keep segments cacheable.
From live to replay
The broadcast should not vanish when it ends. Record the input (record all in nginx-rtmp, or an archive output in MediaLive) to S3, then feed that file into the existing on-demand pipeline. For Northlight, that is the MediaConvert flow in one S3 bucket, one CloudFront: the recording becomes a normal video with a proper ladder, under the same public media domain, and the live playlist URL is swapped for the replay URL once transcoding completes.
Failure modes worth naming
- GOP and segment length disagree. Uneven segments, ABR switching glitches, stalls.
- ffmpeg exits and nobody notices. nginx-rtmp respawns
execprocesses by default, but alert on "stream live, no new segments in 10 seconds" anyway. - Disk fills with segments. Keep
hls_cleanupon (the default) and do not pointhls_pathat a tiny root volume. - Leaked stream key. Someone else broadcasts on your channel. Per-event keys and
on_publishchecks prevent it. - Playlist cached too long. The stream looks frozen for viewers while the encoder looks healthy in the studio.
Takeaways
- Live HLS is ingest, transcode, segment, deliver; decide managed vs self-hosted for the first three, and use CloudFront for the last either way.
- Match the encoder's keyframe interval to your segment length.
- Cache playlists for about a second and segments for a long time.
- Authenticate pushes with short-lived, per-event stream keys.
- MediaLive buys redundancy at a per-hour price; a single EC2 box is cheaper and is a single point of failure.
- Record every broadcast and send it through the on-demand pipeline for replay.