Readstore vs synchronous calls between services
· 5 min read
As soon as a system is split into services, the same question shows up: when service A needs data owned by service B, does it ask for it on the spot or keep a copy?
The default answer is usually to ask: an HTTP call, and if it’s slow, a cache in front. It’s the simplest thing to build and it works well at first. The trouble starts when those calls chain together.
On an e-commerce platform serving more than two million requests a month, one of the decisions that did the most for stability between services was building an internal readstore: a local, read-only copy of the data other services needed to query all the time. This post isn’t about how to implement one. It’s about when it pays off and what you give up in exchange.
The problem: synchronous chains
A user request that crosses several services synchronously has two problems that grow with every hop.
Latency adds up. If the catalogue service calls the pricing service, which calls the stock service, the user waits for all three plus the network in between. The 99th percentile is worse still, because any one of the three having a bad moment is enough.
Availability multiplies down. If each service is up 99.9% of the time, a request that needs all three at once is up 99.7% at best. Every synchronous dependency you add is one more way for your service to fail for reasons that aren’t its fault.
User ──▶ Catalogue ──HTTP──▶ Pricing ──HTTP──▶ Stock
│ │ │
latency = L1 + L2 + L3
availability = A1 × A2 × A3
There’s a third, less visible effect: operational coupling. If Pricing goes down, Catalogue goes down with it. The two teams can no longer deploy, scale or have incidents independently, which was the whole point of splitting them.
The options
1. Synchronous call plus cache
It’s the first reaction, and often it’s enough. The cache lowers average latency and absorbs some of the remote service’s failures.
But it has clear limits. A cache only helps with what it has already seen: a miss still depends on the remote service. TTLs force a choice between fresh data and protection. And when the cache is emptied (a deploy, a restart, a new key), all traffic hits the origin at once, exactly when it can least take it.
2. Event-fed readstore
The service that owns the data publishes an event whenever it changes. The consuming service listens to those events and keeps its own copy, shaped exactly the way its queries need it.
Pricing ──"price updated" event──▶ queue/bus ──▶ projector ──▶ Readstore
│
User ──▶ Catalogue ──local query───────────────────────────────────┘
At request time there’s no remote call anymore: the read is local. If Pricing is down, Catalogue keeps answering with the latest data it knows. And when Pricing comes back, the pending events bring the copy up to date.
3. Shared database
Both services reading from the same tables. It’s the fastest option to build, and I rule it out almost every time: it turns one service’s schema into a public contract nobody can change without coordinating with everyone else. It’s a distributed monolith with extra latency.
Side by side
| Sync + cache | Event-fed readstore | Shared DB | |
|---|---|---|---|
| Read latency | Variable (hits/misses) | Low and stable | Low |
| If the owner is down | Fails on cache miss | Keeps working | Depends on the DB |
| Data freshness | Up to the TTL | Seconds behind | Immediate |
| Build cost | Low | Medium-high | Very low |
| Coupling | Medium (API contract) | Low (event contract) | High (schema) |
The price: what nobody tells you about readstores
Eventual consistency
The copy is always slightly behind the original. Usually by milliseconds or seconds, but if the queue backs up it can be minutes.
The useful question isn’t “is eventual consistency acceptable?” but “on which screens is it acceptable?”. Showing a listing with a price that changed five seconds ago is usually perfectly fine. Charging that price isn’t. My rule: the readstore is for displaying; decisions are made against the data owner.
Maintaining the projection
A readstore isn’t a cache that fills itself. It’s code you have to maintain:
- Replays. If the projection had a bug, or you need a new field, you must be able to rebuild it from scratch. That means being able to re-read events or do an initial load from the owner.
- Event versioning. The event is now a contract between teams. Adding fields is easy; renaming or removing them means living with two versions for a while.
- Ordering and duplicates. Events can arrive twice or out of order. The projection has to be idempotent and know how to discard an event older than the data it already holds.
- Observability. You need to know how far behind the copy is. A stale readstore nobody knows about is worse than a call that fails visibly.
Who owns the data
This is the point most often forgotten. With a readstore, the data still has a single owner. The copy is read-only and nothing but the projector writes to it. If a service starts writing to its copy “because it’s faster”, you end up with two sources of truth drifting apart, and that’s far worse than any latency problem.
When I’d pick it and when I wouldn’t
I’d pick it when:
- One service constantly reads another’s data, and that read sits on the critical path of user requests.
- The data changes far less often than it’s read.
- The screen can live with a few seconds of delay.
- I want a service to keep working even when its dependency is down.
- Event infrastructure already exists, or it makes sense to build it.
I wouldn’t when:
- The read needs the exact value right now (balances, payments, stock reservations at checkout).
- Call volume is low and a cache with a sensible TTL does the job.
- The team doesn’t yet have the operational maturity to run queues, replays and lag alerts. A poorly maintained readstore causes more problems than it solves.
- The services were split for no clear reason, and the real question is whether they should be one.
The underlying idea
A readstore trades latency and availability for consistency and operational complexity. It isn’t a free upgrade: it’s a trade. It pays off when the synchronous dependency is on the critical path and the business can tolerate a few seconds of delay on reads.
Done well, the effect shows up in the hardest thing to achieve in a distributed system: one service failing no longer automatically means the others fail too.