Skip to main content
← Back to articles

Database Sharding: When We Finally Pulled the Trigger

Shard keys, cross-shard queries, and operational overhead after years on one Postgres.

Nestlancer Editorial

Share

Shard keys, cross-shard queries, and operational overhead after years on one Postgres. The team prioritized transactional consistency while migrating on a Nestlancer-style stack—gateway at the edge, domain NestJS services, Prisma on Postgres, async work through RabbitMQ, and media on Backblaze B2.

Where we started

The team inherited single-tenant marketplace data outgrew vertical scaling budgets. The triggering incident was clear: vacuum and index rebuild windows overlapped peak freelancer activity. Leadership funded the refactor when customer-visible latency and support load rose together—not when the diagram looked messy.

Architecture before

  • Single deploy artifact coupling unrelated domains
  • Synchronous cross-module HTTP with partial timeouts
  • Mixed read/write traffic on one database primary
  • Ad-hoc file storage complicating virus scan and CDN caching

Migration timeline

Discovery

Mapped single-tenant marketplace data outgrew vertical scaling budgets and inventoried which routes could move behind the gateway without user-visible changes.

Strangler cutover

Routed new traffic through NestJS handlers while legacy paths drained over 3 weeks.

Reliability hardening

Added Prisma, read replicas, gateway with explicit SLOs owned by a 7-engineer platform squad.

Stabilize and learn

Ran game days for DLQ replay, replica lag failover, and presigned upload expiry edge cases.

Patterns we applied

  • Prisma migrations shipped expand/contract to avoid hard downtime windows.
  • Read replicas via a dedicated Prisma read client served list and analytics traffic.
  • API gateway centralized auth, OpenAPI, and aggregation for mobile clients.

Code sketch from the cutover

@Processor('blog.events')
export class BlogEventsConsumer {
  @EventPattern('post.published')
  async onPublished(@Payload() event: PostPublishedEvent) {
    await this.searchIndexer.index(event.postId);
  }
}

Results

  • sharded tenant archives—primary size flatlined while MAU grew 2.3×
  • On-call pages for queue backlog fell after outbox lag dashboards went live
  • Product teams could ship blog and portfolio changes without redeploying payments
  • Support tickets citing 'stale listings' dropped once read paths moved to replicas

Retrospective checklist

  • Game day DLQ replay documented with ordering notes
  • Replica lag runbook tested in staging monthly
  • Presigned upload TTL aligned with mobile retry policy
  • Gateway error envelope consistent across all domain services
  • Post-incident templates link to dashboards—not screenshots

Lesson

Microservices did not arrive on day one. The team earned splits by proving operational ownership per domain—not by copying a reference diagram. The Nestlancer patterns above were adopted only after metrics justified the coordination cost.

Comments

Loading comments…

Related posts