Next.js App Router SEO for a News Site: A Practical Guide
SEO for a Next.js site: put the story in the HTML, not just the browser
A practical App Router SEO pass: per-page metadata, server-rendered content, structured data, sitemaps and robots, using a fictional newsroom, Northlight Press, as the demo.
Imagine Northlight Press: a digital newsroom with a public reading site built on the Next.js App Router. Stories load fine in a browser. Readers are happy. Search traffic is close to zero.
Open any article and look at the browser tab: every page says "Northlight Press". Share a story on WhatsApp and the preview shows the site name and a generic image. View the page source and the article body is missing; there is a loading skeleton where the story should be.
The site was not broken. It was client-rendered. The page shell came from the server, then React fetched the story in the browser. Humans see the result. Crawlers and link previews mostly see the shell.
This is the pass that fixed it, in the order that mattered.
Demo hostnames for this write-up:
- Production:
northlight.press - Staging:
staging.northlight.press
1. Put the story's own title in the head
The App Router gives every route a generateMetadata function. It runs on the server, so it can fetch the story and describe it before any HTML is sent.
// app/stories/[slug]/page.tsx
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { slug } = await params;
const story = await getStory(slug);
if (!story) return { title: 'Story not found', robots: { index: false } };
const description = summarize(story.excerpt || story.body);
const url = `${SITE_URL}/stories/${story.slug}`;
return {
title: story.title,
description,
alternates: { canonical: url },
openGraph: {
type: 'article',
url,
title: story.title,
description,
images: [story.coverImage ?? '/og-default.jpg'],
publishedTime: story.publishedAt,
modifiedTime: story.updatedAt,
section: story.category?.name,
},
twitter: { card: 'summary_large_image', title: story.title, description },
};
}
In the root layout, set the defaults once:
export const metadata: Metadata = {
metadataBase: new URL(SITE_URL),
title: { default: 'Northlight Press: News and Live Video', template: '%s | Northlight Press' },
description: 'Independent news, live broadcasts and video.',
openGraph: { type: 'website', siteName: 'Northlight Press', images: ['/og-default.jpg'] },
twitter: { card: 'summary_large_image' },
robots: { index: true, follow: true, googleBot: { 'max-image-preview': 'large' } },
};
Now the tab reads "Council approves new transit line | Northlight Press", and a shared link shows that story's headline and cover, served from a private S3 bucket behind CloudFront.
Details that usually bite:
metadataBaseis required for relative image URLs. Without it,og:imagepoints nowhere useful.- Descriptions come from HTML. Strip tags and entities from the body, collapse whitespace, and cut at a word boundary around 160 characters.
- Child
openGraphreplaces the parent's, it does not merge. If a section page setsopenGraph: { title }, it loses the default image and site name. Repeat them. - Do not 404 on an API hiccup. If the fetch fails, return a
noindextitle instead ofnotFound(). A temporary outage should not tell Google the story is gone.
2. Server-render the content, not just the title
A correct <title> on an empty page is half the job. Google does render JavaScript, but later and less reliably than plain HTML, and most link-preview bots never run it.
Northlight's reader component uses React Query. The fix was to fetch the story on the server and seed the cache with it, so the same component renders with data on the first pass:
export default async function StoryPage({ params }: Props) {
const { slug } = await params;
const story = await getStory(slug);
const queryClient = new QueryClient();
if (story) queryClient.setQueryData(['story', 'slug', 'en', slug], story);
return (
<HydrationBoundary state={dehydrate(queryClient)}>
<StoryReader slug={slug} />
</HydrationBoundary>
);
}
No rewrite of the reader, no second rendering path. The headline is an <h1> in the HTML and the full article text is in the first response.
Two things to check:
- The query key must match exactly, including any locale segment. If the client key includes the reader's language, seed the default one; other languages refetch after hydration as before.
fetchwithnext: { revalidate }on the server, sogenerateMetadataand the page share one cached API call instead of two.
3. Tell search engines what the page is
Structured data (JSON-LD) turns "a page with text" into "a news article published at 09:14 by this author in Politics". It feeds rich results and Google News.
export function JsonLd({ data }: { data: object }) {
return (
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(data).replace(/</g, '\\u003c') }}
/>
);
}
The < escape matters: story text is user-authored, and a stray </script> in a headline would otherwise close the tag.
What Northlight emits:
- Every page:
NewsMediaOrganization(name, logo) andWebSitewith aSearchAction, so Google knows the site has a search box at/search?q=. - Story pages:
NewsArticlewith headline, image,datePublished,dateModified, author and section, plus aBreadcrumbList. - Video pages:
VideoObjectwith thumbnail, upload date and an ISO 8601 duration (PT4M12S).
Validate with Google's Rich Results Test before shipping. Invalid structured data is ignored silently.
4. Sitemaps that list every story
app/sitemap.ts returns an array and Next serves it as /sitemap.xml:
export const revalidate = 900;
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const [stories, videos] = await Promise.all([getAllStories(), getAllVideos()]);
return [
{ url: SITE_URL, changeFrequency: 'hourly', priority: 1 },
...stories.map((s) => ({
url: `${SITE_URL}/stories/${s.slug}`,
lastModified: s.updatedAt,
images: s.coverImage ? [s.coverImage] : undefined,
})),
...videos.map((v) => ({
url: `${SITE_URL}/videos/${v.id}`,
videos: [{ title: v.title, description: v.description, thumbnail_loc: v.thumbnail }],
})),
];
}
getAllStories pages through the public API at 100 per request, with a hard cap on pages so a bad pages value from the API can never loop forever.
A newsroom also wants a Google News sitemap: only stories from the last two days, in the news: namespace. Next's sitemap.ts does not emit that format, so it is a small route handler at app/news-sitemap.xml/route.ts that returns XML with Content-Type: application/xml. Escape &, <, > and quotes in titles.
5. robots and noindex, used correctly
app/robots.ts:
export default function robots(): MetadataRoute.Robots {
return {
rules: [{ userAgent: '*', allow: '/', disallow: ['/account', '/sign-in', '/onboarding'] }],
sitemap: [`${SITE_URL}/sitemap.xml`, `${SITE_URL}/news-sitemap.xml`],
};
}
Then mark thin or private pages robots: { index: false } in their metadata: search results, settings, sign-in.
The common mistake: disallowing a page in robots.txt that you want de-indexed. A disallowed page cannot be crawled, so Google never sees its noindex tag. Use disallow to save crawl budget on pages that never mattered, and noindex for pages that should drop out of results.
The newsroom's admin desk gets the opposite of all this: no sitemap, no structured data, a noindex header and a robots.txt that disallows everything. SEO belongs on the reading site only.
6. The small things that add up
- A 1200 × 630 default share image for pages without a cover.
app/manifest.tswith name, colours and icons, so the site can be added to a home screen.themeColorin theviewportexport, so mobile browser chrome matches the brand.- One
SITE_URLsetting per environment. Staging must set its own, or its canonical links and sitemap will point at production, and production's will point at staging if someone copies the wrong env file.
How to check it worked
Before trusting it, look at what a crawler gets:
curl -s https://northlight.press/stories/council-approves-transit-line | grep -o '<title>[^<]*</title>'
curl -s https://northlight.press/sitemap.xml | grep -c '<url>'
curl -s https://northlight.press/robots.txt
Then paste a story URL into the Rich Results Test, and submit both sitemaps in Search Console.
Takeaway
| Problem | Fix in the App Router |
|---|---|
| Every tab shows the site name | generateMetadata per route + title template |
| Crawlers see a loading skeleton | Fetch on the server, seed the client cache |
| Plain text, no rich results | JSON-LD: NewsArticle, VideoObject, WebSite |
| Google finds stories slowly | sitemap.ts + a Google News sitemap route |
| Private pages showing in results | noindex metadata, not just disallow |
The site did not need a rewrite. It needed the server to say, in plain HTML, what each page is about before the browser ever runs a line of JavaScript.