Skip to main content
← Back to articles

Zero-Downtime Auth Service Migration

Session compatibility, token rotation, and dual-write strategies during cutover.

Nestlancer Editorial

Share

Session compatibility, token rotation, and dual-write strategies during cutover. The team prioritized transactional consistency while migrating on a Nestlancer-style stack—gateway at the edge, domain NestJS services, Prisma on Postgres, async work through RabbitMQ, and media on Backblaze B2.

Where we started

The team inherited session tables colocated with user profile data in one auth module. The triggering incident was clear: a schema migration for MFA threatened a hard downtime window. Leadership funded the refactor when customer-visible latency and support load rose together—not when the diagram looked messy.

Architecture before

  • Single deploy artifact coupling unrelated domains
  • Synchronous cross-module HTTP with partial timeouts
  • Mixed read/write traffic on one database primary
  • Ad-hoc file storage complicating virus scan and CDN caching

Migration timeline

Discovery

Mapped session tables colocated with user profile data in one auth module and inventoried which routes could move behind the gateway without user-visible changes.

Strangler cutover

Routed new traffic through NestJS handlers while legacy paths drained over 11 weeks.

Reliability hardening

Added gateway, outbox, Prisma with explicit SLOs owned by a 7-engineer platform squad.

Stabilize and learn

Ran game days for DLQ replay, replica lag failover, and presigned upload expiry edge cases.

Patterns we applied

  • API gateway centralized auth, OpenAPI, and aggregation for mobile clients.
  • Transactional outbox kept Postgres commits and RabbitMQ publishes consistent.
  • Prisma migrations shipped expand/contract to avoid hard downtime windows.

Code sketch from the cutover

@Processor('blog.events')
export class BlogEventsConsumer {
  @EventPattern('post.published')
  async onPublished(@Payload() event: PostPublishedEvent) {
    await this.searchIndexer.index(event.postId);
  }
}

Results

  • dual-wrote sessions for two releases—zero forced logout during cutover
  • On-call pages for queue backlog fell after outbox lag dashboards went live
  • Product teams could ship blog and portfolio changes without redeploying payments
  • Support tickets citing 'stale listings' dropped once read paths moved to replicas

Retrospective checklist

  • Game day DLQ replay documented with ordering notes
  • Replica lag runbook tested in staging monthly
  • Presigned upload TTL aligned with mobile retry policy
  • Gateway error envelope consistent across all domain services
  • Post-incident templates link to dashboards—not screenshots

Lesson

Microservices did not arrive on day one. The team earned splits by proving operational ownership per domain—not by copying a reference diagram. The Nestlancer patterns above were adopted only after metrics justified the coordination cost.

Comments

Loading comments…