Cutting response time by 90% without rewriting the service
· 5 min read
There’s a very common temptation when a service is slow: rewrite it. Another language, another framework, another architecture. It’s almost always a bad idea. What’s slow is rarely the whole service; it’s usually one specific part nobody has measured yet.
This is the story of how a four-engineer team cut the response time of a core service by 90% on an e-commerce platform, without rewriting it. There was no silver bullet. There was a method.
Context
The platform handles more than two million requests a month. The backend is split across several TypeScript services built with NestJS, with MongoDB as the database and Redis available for caching.
One of those services sits at the centre of everything: almost every important page or flow goes through it. When that service is slow, everything else is slow. And it was.
We were a team of four, owning the backend and part of the frontend. At that size there’s no room for a months-long project that freezes all other work.
Constraints
Before touching anything, we set the rules of the game:
| Constraint | Consequence |
|---|---|
| No business downtime | No maintenance windows, no big-bang migrations. |
| No full rewrite | Incremental changes, each one deployable on its own. |
| Small team | Prioritise by impact: the heaviest problem first. |
| Other teams depend on the service | The API contract can’t change. |
These constraints weren’t an obstacle; they were the best guide. They force you to do what you should be doing anyway: measure, go after what hurts most, and ship in small steps.
Diagnosis: measure before you argue
On any team, everyone has a theory about why something is slow. “It’s the database.” “It’s the ORM.” “It’s Node.” Theories are worthless until they’re measured.
The first step was getting real visibility into where each request spends its time. Not the service’s average latency, which hides almost everything, but the breakdown: how much is our own logic, how much is MongoDB queries and how much is calls to other services.
Request to the core service
│
├─ own logic ................ small
├─ MongoDB queries .......... how many? which indexes?
└─ calls to other services
├─ service A ........... sequential or parallel?
├─ service B ........... repeated on every request?
└─ service C ........... what happens when it's slow?
With that breakdown, the usual suspects in a service that has grown for years show up:
- Chained synchronous calls to other services: each one’s latency adds to the next.
- Data that rarely changes but is fetched on every request, from the database or another service.
- Queries without the right index, or fetching far more than is used.
- Repeated work within the same request.
The point of diagnosis isn’t to find one culprit but to rank the culprits by weight. That ranking decides the plan.
Design: attack in order of impact
With causes ranked, the plan almost writes itself. Each change had to meet two conditions: target a measured cause, and be deployable on its own.
1. Stop asking for what we already know
One of the most expensive patterns in a service like this is repeatedly fetching data that belongs to other services and rarely changes. The textbook answer is a Redis cache in front of those calls. It works, but it has limits: the first access is still slow, and if the other service goes down, the cache eventually expires.
So we went a step further and built an internal readstore: a local copy of the data from other services that we need to answer requests, kept up to date when that data changes. The core service no longer asks anyone on the critical path; it reads from its own store.
Before: client ─▶ core service ─▶ service A ─▶ service B ─▶ ...
After: client ─▶ core service ─▶ local readstore
▲
services A, B ───────┘ (update the readstore when their data changes)
That paid off twice: lower latency and more resilience. If another service has a bad day, the core one keeps answering with the latest data it knows.
The price is eventual consistency: for a moment, the readstore may lag behind the source. You have to decide deliberately which data can tolerate that and which can’t. That trade-off deserves its own post.
2. Better communication between services
The calls that remained necessary were reviewed one by one: which could run in parallel instead of in sequence, which had sensible timeouts, and which could degrade gracefully when the other side failed.
3. The basics, done right
The rest was unglamorous but necessary work: MongoDB indexes aligned with the real queries, projections that only fetch the fields in use, and removing duplicated work within a single request.
Outcome
The core service’s response time dropped by 90%. Without a rewrite, without stopping the business and without changing its contract with other teams.
None of it was free:
- More moving parts. The readstore is one more component to watch: you need to know when it falls behind and how to rebuild it.
- Eventual consistency. Some data may take a moment to reflect a change. That’s acceptable where it is, and it has to be documented.
- Measurement discipline. Keeping the gain means measuring continuously; otherwise latency creeps back with every new feature.
What I learned
Measure before you argue. The team’s gut feeling about the cause rarely matches the data completely. The per-request breakdown reshuffled our priorities.
A service’s latency is the sum of its dependencies’. In a distributed system, the fastest service in the world is slow if its critical path waits on three others. Taking dependencies off that path is worth more than optimising code.
A cache and a readstore are not the same thing. A cache speeds up what’s already been requested. A readstore changes who owns the data. The second is more work, but it buys resilience as well as speed.
Constraints help. “No rewrite, no downtime” forced us into small, measurable, reversible changes. Which is exactly how it should always be done.
Rewriting is tempting because it feels like a fresh start. But what’s needed is almost never a new service; it’s knowing exactly where the existing one loses its time.