</>Oriol Martí
← Blog

Anatomy of this site: Astro, S3 and CloudFront for next to nothing

· 6 min read

For years this site lived on shared PHP hosting, uploaded over FTP. It worked, but it had everything I don’t want in a system: a server to maintain, a database to serve text that almost never changes, and manual deployments.

When I rebuilt it, I started from one idea: a content site doesn’t need a server. The cheapest and most resilient architecture is the one with nothing that can go down. This post is the full diagram of where it ended up, why each piece is there and what it costs.

The diagram

                        ┌──────────────┐
  urimarti.com  ──DNS──▶│   Route53    │──▶ A/AAAA alias
                        └──────────────┘          │
                                                  ▼
  Browser ──HTTPS────▶ ┌────────────────────────────────────┐
                       │ CloudFront                         │
                       │  ├─ ACM certificate (us-east-1)    │
                       │  ├─ CloudFront Function (router)   │
                       │  └─ Edge cache                     │
                       └────────────────────────────────────┘
                                      │ OAC (signed)
                                      ▼
                       ┌────────────────────────────────────┐
                       │ Private S3 (eu-west-1)             │
                       │  HTML + assets built by Astro      │
                       └────────────────────────────────────┘

  git push main ──▶ GitHub Actions ──▶ astro build ──▶ cdk deploy

Six pieces, each with a single job:

Piece Job
Astro Generate static HTML at build time.
S3 Store the files. Nothing else.
CloudFront Serve them over HTTPS, cached, close to the user.
CloudFront Function Rewrite and redirect URLs at the edge.
Route53 + ACM DNS and certificate.
CDK + GitHub Actions Make all of the above code that deploys itself.

Astro: zero JavaScript by default

For a content site, Astro has the advantage I care about most: by default it ships no JavaScript to the browser. Every page is HTML and CSS. The little interactivity there is (the theme toggle, the cookie banner) comes from small, independent scripts.

What I ruled out was Next.js. I use it every day and it’s a great tool, but here it would solve problems I don’t have: server rendering, API routes, incremental caching. With output: export it can generate static files too, but it drags along a React runtime this site doesn’t need.

Content is Markdown files in the repo with a schema validated at build time. If a post is missing its description, the build fails. Languages are resolved at build time as well: each page is a single file that generates both /about/ and /en/about/.

S3: a disk, not a web server

S3 has a “static website hosting” mode, but it requires a public bucket and doesn’t serve HTTPS. I ruled it out.

The bucket is private: public access blocked, encrypted and TLS-only. The only way to read it is through CloudFront, which signs every request with Origin Access Control (OAC). If someone finds out the bucket name, there’s nothing they can do with it.

One consequence of this design: with OAC, when a file doesn’t exist S3 doesn’t answer 404 but 403, because from the outside it won’t tell “doesn’t exist” apart from “not allowed”. CloudFront turns both 403 and 404 into the 404.html page with a 404 status.

CloudFront: where everything happens

CloudFront is AWS’s CDN: it terminates HTTPS, caches across its servers worldwide and is the only door into S3. It forces HTTPS and uses the managed CachingOptimized cache policy.

The interesting part is caching, with two policies depending on the file type:

Files Cache-Control Why
/_astro/* (CSS, JS, fonts) max-age=31536000, immutable The name includes a content hash: if it changes, the name changes. Safe to cache for a year.
HTML and everything else max-age=0, s-maxage=3600, must-revalidate The browser always revalidates; CloudFront keeps it for an hour.

Each deploy invalidates /*, so new HTML shows up within seconds instead of waiting out that hour.

The router: a 15-line CloudFront Function

A static site on S3 has a problem: /about/ isn’t a file, /about/index.html is. Something has to translate one URL into the other. That something is a CloudFront Function, JavaScript that runs at the edge before the cache is checked:

function handler(event) {
  var req = event.request;
  var host = req.headers.host.value;
  var uri = req.uri;
  if (host !== 'urimarti.com') {
    return { statusCode: 301, statusDescription: 'Moved Permanently',
      headers: { location: { value: 'https://urimarti.com' + uri } } };
  }
  if (uri.endsWith('/')) {
    req.uri = uri + 'index.html';
  } else if (uri.split('/').pop().indexOf('.') === -1) {
    return { statusCode: 301, statusDescription: 'Moved Permanently',
      headers: { location: { value: uri + '/' } } };
  }
  return req;
}

It does three things: redirects www to the bare domain, adds the trailing slash to URLs missing it, and rewrites /path/ to /path/index.html. That way every page has a single canonical URL, which also matters for SEO.

The alternative was Lambda@Edge, which can do much more (network calls, reading the request body). But it has slower cold starts, slower deploys and higher cost. For rewriting headers and URLs, CloudFront Functions are more than enough.

Route53 and ACM: the domain

The domain is still registered where it always was; only its nameservers point to Route53. The A and AAAA records are aliases to the CloudFront distribution: no IPs to maintain, and the alias works at the zone apex, which a CNAME can’t do.

The HTTPS certificate comes from ACM, free and auto-renewing. There’s one quirk: CloudFront only accepts certificates from us-east-1, even though the rest of the infrastructure lives in Ireland. That’s why the certificate has its own stack in another region. It validates itself, because CDK writes the validation record into the Route53 zone.

Since the domain doesn’t send email, the zone also publishes a null MX, an SPF -all and a DMARC p=reject. Those three records stop anyone from sending spam pretending to be @urimarti.com.

Everything is code: CDK and GitHub Actions

No resource was created by hand in the console, except the DNS zone. Everything lives in a single AWS CDK file with three stacks:

  • Dns: zone records.
  • Cert: the certificate, in us-east-1.
  • Site: bucket, CloudFront, the function, the file upload and the DNS aliases.

GitHub Actions deploys on every push to main: astro build, then cdk deploy. CDK uploads the files to the bucket, deletes the ones that no longer exist and invalidates the CloudFront cache.

The pipeline’s credentials belong to an IAM user that can only assume the roles created by cdk bootstrap. It has no direct permissions on anything. If the key leaked, the possible damage would be limited to what CDK can deploy.

What it costs

Item Monthly cost
Route53 hosted zone $0.50
CloudFront $0 (free tier: 1 TB and 10M requests per month)
CloudFront Functions $0 (2M free invocations per month)
S3 Cents (a few MB)
ACM certificate $0
Invalidations $0 (first 1,000 paths per month are free; /* counts as one)
Total < $1/month

Less than the PHP hosting cost, with a global CDN in front.

What I’d do differently at scale

This architecture is right for a personal site. If it were a product with a team and serious traffic, I’d change a few things:

  • OIDC instead of access keys. GitHub Actions can assume an AWS role with short-lived tokens, without storing any key. For a personal project, scoped keys are enough; in a company, they aren’t.
  • Preview environments. One stack per pull request to review changes before they reach production.
  • Security headers (CSP, HSTS, X-Content-Type-Options) through a CloudFront response headers policy. That’s the next thing I’ll add here.
  • Separate infrastructure from content. Today every push re-evaluates the whole infrastructure, even when only a post changed. CDK leaves untouched what hasn’t changed, but the pipeline takes longer than it needs to.
  • Observability. CloudFront logs and alarms on the 4xx/5xx error rate.

What matters isn’t the list of services but the idea behind it: each piece has a single job, and none of them is a server to maintain. When content rarely changes, building it once and serving it from a cache isn’t an optimisation: it’s the right architecture.