The Agent Migrated Your YAML. It Didn't Migrate Your p99.
The green cutover with a red p99
The migration goes suspiciously well. An agent reads your EKS cluster,
translates every manifest, generates the Terraform for the GKE side, and
opens a change for review. You apply it. Every Deployment reports
Ready, every synthetic check is green, the smoke tests pass.
You cut traffic over.
Then real traffic arrives and p99 latency climbs. Not by a rounding
error — it roughly doubles on a couple of services and stays there. Nothing
is crashing. No pod is CrashLoopBackOff. The dashboards the
agent knew how to reason about are all green. The problem is that a
Kubernetes manifest describes intent, and the agent proved
manifest-equivalence, not behavior-equivalence. Those are different claims,
and the gap between them is exactly the part of a migration that only shows
up under load.
What the agent is genuinely good at
This is not a "don't use the tool" post. Agentic migration tooling — GKE's own EKS-to-GKE assistant, or a home-grown loop wrapped around a capable model — is very good at the mechanical translation that used to eat a sprint:
- Mapping provider-specific annotations to their counterparts
(AWS load balancer annotations to GKE
BackendConfig, IRSA to Workload Identity). - Rewriting StorageClasses, IngressClasses, and CSI provisioner names.
- Generating equivalent IaC and flagging resources with no clean counterpart.
- Cross-referencing two sets of docs faster than any human, and explaining why a given field has no direct equivalent.
That last point matters: a good agent will tell you when a translation is lossy. The failure mode isn't the agent lying — it's that a lossy translation still produces a manifest that applies cleanly and passes a readiness probe. The loss is real; it's just invisible until traffic finds it.
Where manifest-equivalence stops being behavior-equivalence
Storage: same name, different IOPS curve
Suppose your EKS StorageClass pinned explicit performance:
# EKS
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: fast
provisioner: ebs.csi.aws.com
parameters:
type: gp3
iops: "16000"
throughput: "1000" # MB/s, provisioned independently of size
The obvious translation is pd-ssd:
# GKE
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: fast
provisioner: pd.csi.storage.gke.io
parameters:
type: pd-ssd
This applies. It binds. It passes a write smoke test. But on a
Persistent Disk, IOPS and throughput scale with volume size and
instance vCPU count — you don't provision them independently the way gp3
lets you. A 100 GiB pd-ssd does not deliver the same
ceiling as a gp3 volume you'd explicitly dialed to 16,000 IOPS. For a
database or a write-heavy queue, that's a latency regression waiting for the
first busy hour. The behavior-equivalent choice is usually Hyperdisk
(hyperdisk-balanced), where you can set
provisioned-iops-on-create and
provisioned-throughput-on-create — but the agent that
optimized for "smallest valid diff" won't reach for it unless you told it
that the explicit IOPS number was a requirement, not an accident.
Load balancers: the annotations translate, the timeouts don't
The agent will happily map an AWS NLB service to a GKE equivalent and carry your annotations across. What it can't carry across is the default behavior underneath them: idle-connection timeouts, connection-draining semantics, and health-check intervals differ between the two providers' L4/L7 load balancers. Anything holding a long-lived connection — gRPC streams, WebSockets, a database proxy, a message consumer — was implicitly depending on the old defaults. After the move, those connections start getting reset at a different interval, and the symptom is intermittent "connection reset by peer" errors that never correlate with a deploy, because there wasn't one. Pin the timeouts explicitly on both sides so the migration isn't silently changing them:
apiVersion: cloud.google.com/v1
kind: BackendConfig
metadata:
name: long-lived
spec:
timeoutSec: 3600
connectionDraining:
drainingTimeoutSec: 60
Pod density and the network you didn't measure
On EKS with the AWS VPC CNI, each pod gets a real VPC IP and the number of pods per node is bounded by ENI/IP limits for the instance type. GKE is VPC-native with alias IP ranges and a default ceiling around 110 pods per node. If your bin-packing quietly assumed one density and you land on another, you get a different number of nodes for the same workload — which changes noisy-neighbor behavior, per-node network throughput, and how many of your pods share a failure domain. The Deployment spec is identical; the performance envelope is not. This is the kind of thing no manifest diff will ever surface.
Identity: IRSA to Workload Identity
This one the agent usually gets structurally right and semantically incomplete. It knows the annotation changes shape:
# EKS (IRSA)
metadata:
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/app
# GKE (Workload Identity)
metadata:
annotations:
iam.gke.io/gcp-service-account: app@project.iam.gserviceaccount.com
What it can't infer is the IAM binding on the cloud side, the least-
privilege scope that role actually needed, or the fact that some SDK in your
app was reading AWS-specific credential env vars directly. The pod starts.
It's Ready. The first call that actually needs cloud
permissions fails at runtime, well after the health check said everything
was fine.
Build the pipeline so the agent can't hurt you
The announcement language around agentic migration keeps using the word "governance," and it's the right instinct. An agent that can translate and apply is only safe if translation and application are separated by gates it can't skip. The shape that works:
The policy gate is where you encode the things the agent can't be trusted
to remember every time. A Gatekeeper/Policy Controller constraint — or a
Conftest rule in CI — that rejects any workload missing resource requests,
running :latest, or lacking a topology spread constraint costs
almost nothing and catches a whole class of "the translation dropped a
field" bugs before apply:
# conftest / rego: every Deployment must set CPU + memory requests
package main
deny[msg] {
input.kind == "Deployment"
c := input.spec.template.spec.containers[_]
not c.resources.requests.cpu
msg := sprintf("container %q has no CPU request", [c.name])
}
Validate behavior, not readiness
The core mistake is trusting Ready. A readiness probe tells
you a process is listening; it tells you nothing about whether the p99 under
production concurrency matches the cluster you left. So make the last stage
of the pipeline an actual comparison, not a smoke test. Replay a
representative load against both clusters and diff the same PromQL query:
# run identical load (k6/vegeta) at both clusters, then compare:
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)
)
If the target cluster's p99 for a service is meaningfully worse than the source's under the same load, that's your storage/density/timeout gap surfacing — and it surfaces in a staging comparison instead of in an incident channel. The chart below is illustrative, but the shape is the one to watch for: everything looks fine at readiness, and the divergence only appears once you push real concurrency through it.
Lesson
An agent collapses the weeks of mechanical translation into an afternoon, and that's a real win — but it moves the risk rather than removing it. The work that's left is precisely the work that requires knowing your workload's actual performance requirements, which live nowhere in the YAML. Treat the agent as the fastest junior engineer you've ever had: brilliant at the translation, blind to the requirements you never wrote down. Keep the policy gate, the human diff review, and the behavioral validation — and let the agent own the 80% it's genuinely good at, not the 20% that only production can grade.
Hitting something like this in production? I help teams with performance engineering, SRE/observability, and AI-driven root cause analysis — work with me.
Comments
Post a Comment