The Agent Migrated Your YAML. It Didn't Migrate Your p99.

The green cutover with a red p99

The migration goes suspiciously well. An agent reads your EKS cluster, translates every manifest, generates the Terraform for the GKE side, and opens a change for review. You apply it. Every Deployment reports Ready, every synthetic check is green, the smoke tests pass. You cut traffic over.

Then real traffic arrives and p99 latency climbs. Not by a rounding error — it roughly doubles on a couple of services and stays there. Nothing is crashing. No pod is CrashLoopBackOff. The dashboards the agent knew how to reason about are all green. The problem is that a Kubernetes manifest describes intent, and the agent proved manifest-equivalence, not behavior-equivalence. Those are different claims, and the gap between them is exactly the part of a migration that only shows up under load.

What the agent is genuinely good at

This is not a "don't use the tool" post. Agentic migration tooling — GKE's own EKS-to-GKE assistant, or a home-grown loop wrapped around a capable model — is very good at the mechanical translation that used to eat a sprint:

  • Mapping provider-specific annotations to their counterparts (AWS load balancer annotations to GKE BackendConfig, IRSA to Workload Identity).
  • Rewriting StorageClasses, IngressClasses, and CSI provisioner names.
  • Generating equivalent IaC and flagging resources with no clean counterpart.
  • Cross-referencing two sets of docs faster than any human, and explaining why a given field has no direct equivalent.

That last point matters: a good agent will tell you when a translation is lossy. The failure mode isn't the agent lying — it's that a lossy translation still produces a manifest that applies cleanly and passes a readiness probe. The loss is real; it's just invisible until traffic finds it.

Where manifest-equivalence stops being behavior-equivalence

Storage: same name, different IOPS curve

Suppose your EKS StorageClass pinned explicit performance:

# EKS
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: fast
provisioner: ebs.csi.aws.com
parameters:
  type: gp3
  iops: "16000"
  throughput: "1000"   # MB/s, provisioned independently of size

The obvious translation is pd-ssd:

# GKE
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: fast
provisioner: pd.csi.storage.gke.io
parameters:
  type: pd-ssd

This applies. It binds. It passes a write smoke test. But on a Persistent Disk, IOPS and throughput scale with volume size and instance vCPU count — you don't provision them independently the way gp3 lets you. A 100 GiB pd-ssd does not deliver the same ceiling as a gp3 volume you'd explicitly dialed to 16,000 IOPS. For a database or a write-heavy queue, that's a latency regression waiting for the first busy hour. The behavior-equivalent choice is usually Hyperdisk (hyperdisk-balanced), where you can set provisioned-iops-on-create and provisioned-throughput-on-create — but the agent that optimized for "smallest valid diff" won't reach for it unless you told it that the explicit IOPS number was a requirement, not an accident.

Load balancers: the annotations translate, the timeouts don't

The agent will happily map an AWS NLB service to a GKE equivalent and carry your annotations across. What it can't carry across is the default behavior underneath them: idle-connection timeouts, connection-draining semantics, and health-check intervals differ between the two providers' L4/L7 load balancers. Anything holding a long-lived connection — gRPC streams, WebSockets, a database proxy, a message consumer — was implicitly depending on the old defaults. After the move, those connections start getting reset at a different interval, and the symptom is intermittent "connection reset by peer" errors that never correlate with a deploy, because there wasn't one. Pin the timeouts explicitly on both sides so the migration isn't silently changing them:

apiVersion: cloud.google.com/v1
kind: BackendConfig
metadata:
  name: long-lived
spec:
  timeoutSec: 3600
  connectionDraining:
    drainingTimeoutSec: 60

Pod density and the network you didn't measure

On EKS with the AWS VPC CNI, each pod gets a real VPC IP and the number of pods per node is bounded by ENI/IP limits for the instance type. GKE is VPC-native with alias IP ranges and a default ceiling around 110 pods per node. If your bin-packing quietly assumed one density and you land on another, you get a different number of nodes for the same workload — which changes noisy-neighbor behavior, per-node network throughput, and how many of your pods share a failure domain. The Deployment spec is identical; the performance envelope is not. This is the kind of thing no manifest diff will ever surface.

Identity: IRSA to Workload Identity

This one the agent usually gets structurally right and semantically incomplete. It knows the annotation changes shape:

# EKS (IRSA)
metadata:
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/app

# GKE (Workload Identity)
metadata:
  annotations:
    iam.gke.io/gcp-service-account: app@project.iam.gserviceaccount.com

What it can't infer is the IAM binding on the cloud side, the least- privilege scope that role actually needed, or the fact that some SDK in your app was reading AWS-specific credential env vars directly. The pod starts. It's Ready. The first call that actually needs cloud permissions fails at runtime, well after the health check said everything was fine.

Build the pipeline so the agent can't hurt you

The announcement language around agentic migration keeps using the word "governance," and it's the right instinct. An agent that can translate and apply is only safe if translation and application are separated by gates it can't skip. The shape that works:

Discover read EKS state Translate agent → diff Policy gate OPA / Policy Human review diff Apply staged Validate: load test, p99 parity compare vs source cluster regressions feed back

The policy gate is where you encode the things the agent can't be trusted to remember every time. A Gatekeeper/Policy Controller constraint — or a Conftest rule in CI — that rejects any workload missing resource requests, running :latest, or lacking a topology spread constraint costs almost nothing and catches a whole class of "the translation dropped a field" bugs before apply:

# conftest / rego: every Deployment must set CPU + memory requests
package main

deny[msg] {
  input.kind == "Deployment"
  c := input.spec.template.spec.containers[_]
  not c.resources.requests.cpu
  msg := sprintf("container %q has no CPU request", [c.name])
}

Validate behavior, not readiness

The core mistake is trusting Ready. A readiness probe tells you a process is listening; it tells you nothing about whether the p99 under production concurrency matches the cluster you left. So make the last stage of the pipeline an actual comparison, not a smoke test. Replay a representative load against both clusters and diff the same PromQL query:

# run identical load (k6/vegeta) at both clusters, then compare:
histogram_quantile(0.99,
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)
)

If the target cluster's p99 for a service is meaningfully worse than the source's under the same load, that's your storage/density/timeout gap surfacing — and it surfaces in a staging comparison instead of in an incident channel. The chart below is illustrative, but the shape is the one to watch for: everything looks fine at readiness, and the divergence only appears once you push real concurrency through it.

p99 latency (illustrative) source baseline migrated ~2x 0

Lesson

An agent collapses the weeks of mechanical translation into an afternoon, and that's a real win — but it moves the risk rather than removing it. The work that's left is precisely the work that requires knowing your workload's actual performance requirements, which live nowhere in the YAML. Treat the agent as the fastest junior engineer you've ever had: brilliant at the translation, blind to the requirements you never wrote down. Keep the policy gate, the human diff review, and the behavioral validation — and let the agent own the 80% it's genuinely good at, not the 20% that only production can grade.


Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.

Comments

Popular posts from this blog

Performance Testing 102: Little's Law and It's usage in Performance Testing

Performance Testing 104: Workload Modelling Designing & Process

Mastering the Art of Scaling in SaaS Applications