Tutorials
From Markdown to a Live Site with Astro Content Collections
Turn a folder of Markdown into a typed collection where a malformed post fails the build — including the draft flag that has to be honored in three places to stay off the live site.
On this page
The problem, in one folder
Open the folder of Markdown posts behind any blog that’s grown past its first month, and you’ll usually find the same small mess: one post has category: Tutorials, another has category: tutorials, a third omits the field entirely. Nothing about writing Markdown stops you from doing this — a frontmatter block is only a convention until something enforces it. Left alone, that drift shows up as a blank category badge on a live page, or a listing page that quietly renders one fewer card than it should, and you usually find it by noticing something’s missing, not by being told.
This piece is the fix for that, end to end: how a folder of Markdown becomes a typed, validated collection, so a malformed post fails the build instead of reaching a reader — and, because “finished” and “safe to publish” turn out to be two different states, how a draft stays off the live site until you want it there.
The collection, minimal
A content collection starts as two small pieces: something that finds your files, and something that checks their shape. In Astro, that’s a loader and a schema, both declared in one config file (conventionally src/content.config.ts):
import { defineCollection, z } from "astro:content";
import { glob } from "astro/loaders";
const blog = defineCollection({
loader: glob({ pattern: "**/*.md", base: "./src/content/blog" }),
schema: z.object({
title: z.string(),
summary: z.string(),
publishDate: z.string(),
}),
});
export const collections = { blog };That’s the whole shape our own collection is built on: a glob over every .md file under one base folder, and a schema describing what a valid post looks like. Three fields is deliberately the smallest useful version — enough to prove the mechanism before you trust it with anything bigger.
So prove it. Delete title from one post’s frontmatter and run the build. (If you’ve not run one from a terminal before, start here.) It won’t quietly skip that post — it stops, and it tells you which file and which field: a required string is missing where the schema expected one. That failure is the entire point of a schema. A folder of Markdown with no schema fails silently, later, in a browser. A folder validated at build time fails loudly, immediately, in your terminal, with a file name attached.
A schema checks shape, not content. It confirms publishDate is a string; it has no opinion on whether that string is a real, sensible date. A typo that still parses gets through. That’s not a gap you close by writing a better schema — it’s the boundary of what this kind of check can do at all.
Closed sets and the single source of truth
Our actual collection has ten fields, not three: title, summary, category, format, publishDate, featured, draft, and three cross-link fields — prerequisites, related, and usedIn — each a list of other posts’ slugs. featured and draft are booleans that default to false; the cross-link fields default to an empty list, so a post that links nowhere doesn’t have to say so.
summary deserves its own sentence: it’s one field doing four jobs at once — the deck under the headline, the excerpt on the card, the <meta description> tag, and the Open Graph description social platforms read. Write it once, well, and it does all four; write it as an afterthought and all four are worse.
The more interesting design choice is category. It’s not a free-text string — it’s an enum of exactly four slugs — build-logs, tutorials, primers, toolbox — and a post with an unrecognized fifth one fails the build, the same as a missing title. The stored value is the slug; the display label lives somewhere else, in one place, a separate config file, so renaming a category is a one-line change instead of a search-and-replace across every post that uses it.
The rule underneath both choices is the same one: close a field into a fixed set when a typo in it would silently produce a wrong page — a category nobody can navigate to, a broken link nobody notices. Leave it open, a plain string, when a typo only makes one post a little worse, not invisible. title stays a string because a typo there is a bad headline, not a bad build. category became an enum because a typo there used to be worse than that.
Why .md, not .mdx
.md keeps every article page zero-JS by default: no component runtime ships to render a blog post. The two things a rich post usually wants, code blocks and callout boxes, still exist; they’re produced by build-time transforms of ordinary Markdown syntax, not by importing interactive components into the content itself.
The cost is real: you give up dropping a live component into an article. No embedded calculator, no interactive demo, nothing that needs its own JavaScript running in the reader’s browser. If that’s a page you need, .mdx is the honest answer for that one page, not a downgrade from .md, just a different tool for a different job. For a text-first blog, .md was the right default; it isn’t the only correct one everywhere.
The draft flag, and why one flag isn’t enough
draft looks like the simplest field in the whole schema — one boolean, defaulting to false — and it’s tempting to think that setting it is the whole job. It isn’t. A flag only does something if every piece of code that reads posts actually checks it, and “every piece of code” turns out to mean three separate places, not one.
Layer one: the listings. Every surface that lists posts — the homepage, the category pages, the archive — pulls from one shared helper, and that helper excludes anything with draft: true, on every environment, no exceptions. This is the layer people picture when they hear “draft flag,” and it’s necessary but not sufficient.
Layer two: the route itself. A draft still needs to be reachable at the URL the published post will use, before you publish it — that’s how you preview it, share a link with someone, catch a bad heading before the world does. So a second helper makes drafts routable, but only when the build is not running as production. On staging, a draft’s page loads. In production, the same URL doesn’t exist. Which of those two you get depends on reading one environment variable with a fallback, and the read is written to fail closed: if the value is missing or something unrecognized, it’s treated as not production, on purpose, rather than guessing yes.
Layer three: the cross-links. A post can link to another by slug (think “read this next”) and that resolver draws a distinction most flags miss. Link to a slug that doesn’t exist at all, and the build throws and stops you, every time, everywhere. Link to a slug that exists but is still a draft, and it’s dropped silently instead: no broken link on the page, only one fewer suggestion than the author wrote. Same mistake, deliberately different behavior, because an unknown slug is an author’s typo and a draft target is a timing problem that resolves itself once the other post ships.
Here’s the bug that proved three layers were the actual requirement, not two. One component on the site called the raw collection lookup directly, instead of going through the shared helper. It worked, right up until a published post linked to a post that was still a draft: the draft’s card rendered anyway, on the production listing, because the one component that skipped the helper never checked the flag at all. Click it, and you landed on a 404. Nothing about the first two layers caught this — they were both doing their job correctly. The hole wasn’t a fourth layer — it was layer one, at the one call site that never went through the helper. The listing rule was right; nothing said every listing had to go through the thing that enforced it.
A check that scans instead of testing
That bug is also why the build now runs a check before it runs anything else — and it’s worth being precise about what kind of check it is, because it isn’t a test. A test runs code and checks the output. This walks the source and checks who’s allowed to call what.
It has three parts. The first reads every file in the project and finds every call site that touches the raw collection API; anything outside one narrow, named allowlist fails the build on the spot: that’s how the bug above gets caught going forward, not only in the one case that already happened. The second re-checks the allowlisted helpers themselves, confirming the guard logic is still there and hasn’t been quietly edited away. The third is the one to steal: the check also asserts that it found files and call sites at all. Without that, a scanner with a broken path, pointed at an empty folder, would report a clean pass every time — a scanner that walks nothing passes everything, silently, which is a worse failure than not scanning in the first place.
It’s wired straight into the build script, so it runs before the site itself builds; a failing scan is a failing deploy, not a warning someone can scroll past. It oversells easily, so: it proves that every consumer goes through the guard. It does not prove the guard’s own logic is correct. Those are two different claims, and this check only makes the first one — you’d still need a different kind of check, run against actual output, to make the second.
From a file to production
Once a post’s frontmatter and content are ready, the file goes onto the staging branch with draft: true — and it’s already live at the address it will keep once published, immediately, exactly as designed: reachable for review, absent from every listing. Flip that one field to false and push again, and the difference shows up as a difference between two builds: the route that staging already had now appears in the production build too, where before it simply hadn’t been generated. That’s not a guess about what should happen; it’s what changing the flag actually produced.
Getting that change from staging to the live site is a human decision, not an automatic one — a pull request, from a known-good point, that a person reads and merges. The deploy pipeline itself already has its own piece; the short version here is enough: promotion is a gate, and the gate is a person.
Worth checking afterward is something more than “does the page load.” The last time we flipped a post live, the sitemap went from 12 entries to 13, and that post’s own category tile went from 2 to 3 — both snapshots from that one moment, not running totals, because they move every time a post is added, and the four category tiles between them always sum to the size of the whole collection. A new post changes more than its own URL, and the surfaces that summarize the whole collection are worth a glance too.
Copy, watch, and the close
Copy: the collection config and its loader, the rule for when a field becomes a closed enum instead of a free string, the three-layer draft pattern end to end, the scanning check — especially its anti-vacuity arm — and wiring that check into your build script so a failure there is a failure to deploy.
Watch, don’t copy: our specific branch-and-pull-request arrangement, which belongs to the deploy piece this one leans on, and the exact counts above, the field list included, which were true at one moment and will already have moved by the time you read this.
This post went through the same pipeline described above: Markdown file, schema check, draft flag, a human’s merge.