Elasticsearch Migration for Search Features
Reindex strategies, zero-downtime aliases, and relevance tuning.
Nestlancer Editorial

Reindex strategies, zero-downtime aliases, and relevance tuning. The team prioritized transactional consistency while migrating on a Nestlancer-style stack—gateway at the edge, domain NestJS services, Prisma on Postgres, async work through RabbitMQ, and media on Backblaze B2.
Where we started
The team inherited SQL ILIKE search timing out on large freelancer catalogs. The triggering incident was clear: marketing promised faceted search before holiday traffic. Leadership funded the refactor when customer-visible latency and support load rose together—not when the diagram looked messy.
Architecture before
- Single deploy artifact coupling unrelated domains
- Synchronous cross-module HTTP with partial timeouts
- Mixed read/write traffic on one database primary
- Ad-hoc file storage complicating virus scan and CDN caching
Migration timeline
Discovery
Mapped SQL ILIKE search timing out on large freelancer catalogs and inventoried which routes could move behind the gateway without user-visible changes.
Strangler cutover
Routed new traffic through NestJS handlers while legacy paths drained over 12 weeks.
Reliability hardening
Added RabbitMQ, outbox, read replicas with explicit SLOs owned by a 7-engineer platform squad.
Stabilize and learn
Ran game days for DLQ replay, replica lag failover, and presigned upload expiry edge cases.
Patterns we applied
- RabbitMQ workers absorbed notification and indexing fanout with DLQ replay runbooks.
- Transactional outbox kept Postgres commits and RabbitMQ publishes consistent.
- Read replicas via a dedicated Prisma read client served list and analytics traffic.
Code sketch from the cutover
@Processor('blog.events')
export class BlogEventsConsumer {
@EventPattern('post.published')
async onPublished(@Payload() event: PostPublishedEvent) {
await this.searchIndexer.index(event.postId);
}
}
Results
- indexed via outbox events—P95 search dropped from 2.4s to 180ms
- On-call pages for queue backlog fell after outbox lag dashboards went live
- Product teams could ship blog and portfolio changes without redeploying payments
- Support tickets citing 'stale listings' dropped once read paths moved to replicas
Retrospective checklist
- Game day DLQ replay documented with ordering notes
- Replica lag runbook tested in staging monthly
- Presigned upload TTL aligned with mobile retry policy
- Gateway error envelope consistent across all domain services
- Post-incident templates link to dashboards—not screenshots
Lesson
Microservices did not arrive on day one. The team earned splits by proving operational ownership per domain—not by copying a reference diagram. The Nestlancer patterns above were adopted only after metrics justified the coordination cost.
Comments
Loading comments…
Related posts

Case Studies
Cutting Deploy Time from 45 Minutes to Five
CI caching, smaller artifacts, and service-level pipelines after monolith split.

Case Studies
Scaling a Freelance Marketplace Architecture
Matching algorithms, escrow flows, and dispute resolution at growing GMV.

Case Studies
GDPR Compliance Platform Rebuild
Data maps, deletion workflows, and consent logging across microservices.

Case Studies
Migrating from WebSockets to SSE
Simpler infra, CDN friendliness, and trade-offs for one-way realtime feeds.