Anatomy of a serverless SaaS on AWS
· 6 min read
Build the Drop is my SaaS: it syncs catalogues and inventory between AliExpress and Shopify. I build it alone. That shapes every architectural decision more than any technical requirement does: there’s nobody else to pick up an alert at three in the morning.
This post isn’t about specific services. It’s about the shape of the system and the reasoning behind it. It’s the kind of architecture I’d pick today for any product that has to move a lot of data between other people’s APIs, on a small budget, with a team of one.
Diagram
Shopify ──webhooks──▶ ┌──────────────┐
│ API Gateway │──▶ "ingest" function
Dashboard (user) ────▶│ + auth │ │ validates and enqueues
└──────────────┘ ▼
┌──────────────┐
Scheduler ──── every N minutes ───────▶│ Queue (SQS) │◀── retries / DLQ
└──────────────┘
│ batches
▼
┌───────────────────────────┐
│ "sync" functions │
│ ├─ read from AliExpress │
│ ├─ compute the change │
│ └─ write to Shopify │
└───────────────────────────┘
│
▼
┌───────────────────────────┐
│ Database │
│ state of each product │
│ + last applied sync │
└───────────────────────────┘
Four ideas hold the diagram together:
| Idea | What it solves |
|---|---|
| Nothing is processed in the request | The entry point only validates and enqueues. Answering a webhook fast is what stops Shopify from retrying or disabling it. |
| The queue is the centre | It absorbs spikes, enables retries, and separates “what needs doing” from “when I’m allowed to do it”. |
| Idempotent jobs | Running the same sync twice leaves the same result. |
| State lives in the database, not in the function | Any function can die halfway through a batch without losing anything. |
Why serverless for a solo founder
When you’re one person, the scarce resource isn’t CPU, it’s your attention. Every server is a to-do list that adds nothing to the product: security patches, disks filling up, hung processes, manual scaling when a big customer shows up.
With functions and managed services, that list shrinks to almost nothing. In exchange you accept three costs, and it’s worth being clear about them from day one:
- Execution limits. A function can’t work for hours. It forces you to split work into small pieces, which in this kind of system is exactly what you want anyway.
- Cold starts. Irrelevant for background jobs. Barely noticeable in the user dashboard if functions stay small.
- Harder debugging. There’s no server to log into and poke around. Without structured logs and per-job tracing, you’re flying blind.
The other argument is economic: cost scales with usage, not with time. A young SaaS spends many hours with no traffic at all. Paying for servers that are up during those hours is paying to wait.
Bulk sync: queues, batches and third-party API limits
The core problem of Build the Drop isn’t technical in the classic sense: you depend on two APIs you don’t control, each with its own rate limits, outages and formats. The architecture exists mainly to live with that.
API limits are in charge
Shopify and AliExpress cap how many requests you can make per unit of time. If a user with thousands of products triggers a full sync, you can’t do it all at once. And with many users, their limits are per store, but your infrastructure is shared.
So the scheduler doesn’t do the work: it puts messages on the queue, and consumers process them at whatever pace the APIs allow. When an API answers “too many requests”, the message goes back to the queue with a delay instead of failing. The queue turns an external limit into a plain delay.
Batches, not single products
Processing one product at a time multiplies API calls and function starts. Processing the whole catalogue in one run hits the time limits. The middle ground is fixed-size batches: big enough to amortise each run, small enough to finish with headroom and be retried whole without drama.
Idempotency: a sync you can repeat
In a system with queues and retries, every message will be processed more than once at some point. The question isn’t whether it’ll happen, but what happens when it does.
My rule is that a sync job doesn’t say “subtract 3 from stock”, it says “stock must be 12”. You compute the desired state from the source, compare it with the last applied state, and only write if there’s a difference. Running it ten times gives the same result as running it once.
source (AliExpress) ──▶ desired state ──┐
├─▶ different? ──yes──▶ write to Shopify
last applied state (DB) ────────────────┘ │ │
no ▼
▼ save new state
nothing to do
Comparing before writing also saves most write calls: in a large catalogue, only a small share of products changes between passes.
What fails gets set aside
Some messages will never succeed: a product deleted at the source, a revoked user token, data in an unexpected format. Retrying them forever just burns money and floods the logs. After a few attempts they go to a dead-letter queue (DLQ), which I review and can replay once I’ve fixed the cause.
Real costs
I won’t give numbers here: they depend on the number of stores, catalogue sizes and sync frequency, and I’d rather publish them once I have data worth showing.
What I can explain is what drives cost in an architecture like this:
| Item | Driven by |
|---|---|
| Function invocations | Number of batches × duration of each |
| Queue messages | Number of batches and retries |
| Database | Per-product state reads and writes |
| API Gateway | Webhooks and dashboard usage |
| Data transfer | Negligible: small JSON payloads |
The practical consequence is that cost grows roughly in a straight line with customers, and is close to zero at rest. For a young SaaS, that means you know your margin per customer from day one. It also means the biggest cost risk isn’t traffic, it’s a bug retrying in a loop. That’s why billing alarms and retry limits go in from the start, not “when needed”.
What doesn’t scale (yet)
This architecture is meant to grow for quite a while without changes, but I know where it’ll break first:
- Customers with huge catalogues. A full sync of a very large catalogue is still slow, because the bottleneck is the source API’s limit, not my infrastructure. The fix is to sync only what changed, where the source allows it, instead of walking everything.
- Fairness between customers. With a shared queue, one big customer can delay the small ones. The next step is splitting work per customer or by priority.
- Per-customer observability. Today I know whether the system works. I want to know, without digging through logs, when each store last synced successfully and why the last failure failed.
- Me. The real limit of a one-person SaaS isn’t technical. Every decision in this post has the same goal: a system that needs me as little as possible.
The lesson I take away: in a system that depends on other people’s APIs, good architecture isn’t the one that’s fast, it’s the one that assumes everything will fail and still reaches the right state. Queues, batches and idempotency aren’t optimisations. They’re what lets you sleep.