Skip to main content
← Back to articles

Kubernetes Migration Cost Analysis

Real infra bills, engineer time, and when managed services beat self-hosted clusters.

Nestlancer Editorial

Share

Real infra bills, engineer time, and when managed services beat self-hosted clusters. The team prioritized incremental cutover while migrating on a Nestlancer-style stack—gateway at the edge, domain NestJS services, Prisma on Postgres, async work through RabbitMQ, and media on Backblaze B2.

Where we started

The team inherited Compose-based staging that did not reflect production autoscaling behavior. The triggering incident was clear: Black Friday forced manual VM scaling with opaque costs. Leadership funded the refactor when customer-visible latency and support load rose together—not when the diagram looked messy.

Architecture before

  • Single deploy artifact coupling unrelated domains
  • Synchronous cross-module HTTP with partial timeouts
  • Mixed read/write traffic on one database primary
  • Ad-hoc file storage complicating virus scan and CDN caching

Migration timeline

Discovery

Mapped Compose-based staging that did not reflect production autoscaling behavior and inventoried which routes could move behind the gateway without user-visible changes.

Strangler cutover

Routed new traffic through NestJS handlers while legacy paths drained over 11 weeks.

Reliability hardening

Added gateway, read replicas with explicit SLOs owned by a 4-engineer platform squad.

Stabilize and learn

Ran game days for DLQ replay, replica lag failover, and presigned upload expiry edge cases.

Patterns we applied

  • API gateway centralized auth, OpenAPI, and aggregation for mobile clients.
  • Read replicas via a dedicated Prisma read client served list and analytics traffic.

Code sketch from the cutover

await this.prisma.$transaction(async (tx) => {
  const post = await tx.post.update({ where: { id }, data: { status: 'PUBLISHED' } });
  await tx.outboxEvent.create({
    data: { aggregateId: post.id, type: 'post.published', payload: { postId: post.id } },
  });
});

Results

  • rightsized K8s requests saved 19% monthly spend with stable P95
  • On-call pages for queue backlog fell after outbox lag dashboards went live
  • Product teams could ship blog and portfolio changes without redeploying payments
  • Support tickets citing 'stale listings' dropped once read paths moved to replicas

Retrospective checklist

  • Game day DLQ replay documented with ordering notes
  • Replica lag runbook tested in staging monthly
  • Presigned upload TTL aligned with mobile retry policy
  • Gateway error envelope consistent across all domain services
  • Post-incident templates link to dashboards—not screenshots

Lesson

Microservices did not arrive on day one. The team earned splits by proving operational ownership per domain—not by copying a reference diagram. The Nestlancer patterns above were adopted only after metrics justified the coordination cost.

Comments

Loading comments…

Related posts