Key takeaways

Short answer: Autoscaling and optimization are different problems.

  • Fleet-wide CPU utilization averages just 8%, according to the Cast AI 2026 State of Kubernetes report across tens of thousands of production clusters.
  • CPU overprovisioning reached 69% in 2026, up from 40% the prior year; memory overprovisioning hit 79%.
  • HPA scales replica count only. It never touches resource requests. Karpenter provisions against those same inflated numbers.
  • In a 7-day adversarial EKS benchmark, full Cast AI optimization cut weekly spend by 43% versus Karpenter-native alone.
  • OOM kills dropped from 40-50 per 30-minute monitoring window to near zero after automated rightsizing, while provisioned CPUs fell by roughly 50%.
  • 88% of organizations report year-over-year TCO increases (Spectro Cloud 2025 State of Kubernetes report); 42% cite cost as their top Kubernetes challenge (Spectro Cloud 2025 State of Kubernetes report).

What is the autoscaling gap?

The autoscaling gap is the difference between what your cluster provisions and what your workloads actually consume. Kubernetes autoscalers, including HPA, VPA, Karpenter, and Cluster Autoscaler, all make decisions based on declared resource requests. None of them can see real consumption unless you close that feedback loop deliberately. The result: clusters that scale correctly on paper but waste most of their compute budget in practice. The Cast AI 2026 report, covering tens of thousands of clusters, found average CPU utilization at 8%. The other 92% is cost with no return. An autonomous cluster closes that gap: it continuously adjusts both pod resource requests and node capacity based on actual consumption data, without manual tuning cycles or incident-driven changes. The patterns apply across EKS, GKE, and AKS, though the benchmark data is EKS-based and Spot pricing, interruption rates, and instance selection behavior differ by cloud provider.

This is not a misconfiguration problem. It is a structural one. The tools engineers rely on were not designed to optimize; they were designed to prevent failures. Understanding where each tool stops is the first step to closing the gap. An autonomous cluster closes that gap: it continuously adjusts both pod resource requests and node capacity based on actual consumption data, without manual tuning cycles or incident-driven changes.

The Two-Layer Autoscaling Stack

Most Kubernetes autoscaling discussions conflate two fundamentally different problems. Pod autoscaling handles replica count and resource sizing. Node autoscaling handles cluster capacity. They operate on different signals, different timescales, and different primitives. Treating them as interchangeable is what creates the gap.

The Two-Layer Autoscaling Stack separates these concerns clearly. Each layer has distinct responsibilities, distinct failure modes, and distinct optimization levers. Fixing one layer without addressing the other leaves most of the waste intact.

Layer 1: Pod autoscaling (HPA and VPA)

HPA (Horizontal Pod Autoscaler) watches a metric, usually CPU or memory utilization as a percentage of the request, and adjusts replica count. When load increases, you get more pods. When load drops, pods scale down. Simple, battle-tested, widely used.

HPA has one critical limitation: it does not change resource requests. If a pod declares requests.cpu: 500m and actually uses 40m, HPA will scale to 10 replicas during a traffic spike and back to 2 at night. The 500m request per replica stays untouched throughout. Every scheduling and bin-packing decision downstream operates on that inflated number.

VPA (Vertical Pod Autoscaler) fixes the request. It watches real consumption and recommends updated CPU and memory values. In theory, it closes the gap that HPA creates. In practice, two constraints limit it:

  • VPA requires a pod restart to apply new requests, disrupting running workloads. In Kubernetes 1.33 and later, in-place pod vertical scaling (KEP-1287) allows memory and CPU updates without a pod restart. Production adoption is still limited by controller compatibility and the HPA conflict remains when both target the same resource metric. Note: EKS, GKE, and AKS require explicit InPlacePodVerticalScaling=true feature gate enablement below Kubernetes 1.35. Most production managed clusters do not have this gate enabled by default as of mid-2026.
  • VPA cannot run on the same metric as HPA for the same workload; doing so creates conflicting recommendations.

Most teams deploy HPA broadly and restrict VPA to batch jobs or development environments. Production pods run with static, inflated resource requests indefinitely. That is the default state, not an edge case.

Layer 2: Node autoscaling (Karpenter and Cluster Autoscaler)

Karpenter and Cluster Autoscaler (CAS) work at the node level. When pods cannot be scheduled, they provision new nodes. When nodes look underutilized, they consolidate and terminate.

Karpenter provisions nodes in 45-60 seconds (with AMI pre-caching enabled; cold provisioning without caching takes 2-4 minutes, similar to CAS). CAS takes 3-4 minutes. That speed difference matters during traffic spikes. For scale-down, Karpenter”s consolidation logic is more aggressive than CAS”s default 50% threshold, which compares requested resources to node capacity, not actual consumption.

Both tools share the same structural problem: they see declared requests, not real usage. If your pods request 500m CPU and use 40m, the node autoscaler provisions and schedules as if every pod will need 500m. Better consolidation logic helps at the margins. It does not fix the underlying signal.

Why engineers inflate resource requests

This is not ignorance. Engineers inflate requests for rational reasons.

The primary driver is OOM kills. When a container exceeds its memory limit, Kubernetes terminates it. OOM kills page your on-call team at 2 AM. The incentive to set limits high is immediate and visceral. The cost of overprovisioning is invisible, diffused across next month”s cloud bill.

CPU throttling creates similar pressure. When a container exceeds its CPU request, the kernel throttles it. Latency spikes. SLOs break. Engineers add headroom. No alert fires for overprovisioning.

There is also no feedback loop. Kubernetes does not tell you that your pod requested 500m and used 40m. You need metrics-server, Prometheus, custom dashboards, and organizational discipline to act on what you find. Most teams have the metrics. Few act on them systematically, because acting means touching production configs and accepting the risk of getting it wrong.

The outcome is predictable. CPU overprovisioning reached 69% fleet-wide in 2026, up from 40% the prior year. Memory overprovisioning hit 79%. These numbers are not outliers from poorly-managed clusters. They are the default state of any cluster where resource requests are a manual, incident-driven process.

What autoscalers actually see: the request vs. reality gap

Your autoscalers operate on a model of your cluster that diverges from reality by a wide margin. The table below makes that concrete.

Signal What autoscalers see What is actually happening
CPU usage Pod requests 500m CPU Pod uses ~40m CPU on average
Memory usage Pod requests 512Mi memory Pod uses ~110Mi at peak
Node utilization Node at 70% of requested capacity Node at 8% of actual compute capacity
CAS scale-down trigger Requested resources below 50% of node capacity Actual usage may be 5-10% of node capacity
HPA scaling signal CPU utilization 8% (40m / 500m request) Actual CPU pressure is near zero
Karpenter node selection Selects node type based on total request sum Selected node is far larger than actual workload needs

Every row is a compounding error. HPA sees low utilization relative to an inflated request and leaves replica count alone. Karpenter picks a node based on the sum of inflated requests. CAS holds that node because requested capacity sits above its 50% threshold. The cluster appears healthy. Monitoring shows “normal” utilization. You are paying for 10-12x what you use, and every tool in your stack confirms everything is fine.

The Karpenter ceiling: where consolidation breaks down

Karpenter is genuinely good at what it does. Faster provisioning, flexible node selection, and better consolidation than CAS make it the right default for most teams. But it has a ceiling, and that ceiling becomes visible under controlled benchmark conditions.

Cast AI ran a 7-day adversarial test on EKS, comparing four configurations with identical workloads and traffic patterns. “Adversarial” means the workload resisted consolidation: variable traffic, mixed pod sizes, and real scheduling constraints. Cast AI engineering disabled workload rightsizing to isolate the consolidation effect. The results:

Configuration Weekly cost Savings vs. Karpenter native
Karpenter native $703.08 Baseline
Karpenter + Cast AI Evictor $639.45 9.1%
Karpenter + Evictor + Karpenter Enterprise Consolidation and Rebalancing (CRB) $591.98 15.8%
Full Cast AI (all layers) $400.83 43%

The 43% figure is a conservative lower bound. Because rightsizing was disabled to isolate consolidation, actual savings with all Cast AI layers enabled exceed this number.

Karpenter consolidation alone delivers nothing here beyond the baseline, because the baseline already uses Karpenter. Each additional layer closes the gap incrementally. But the step from +CRB (15.8%) to full Cast AI (43%) is not incremental; it is a structural shift. That jump reflects what happens when you fix the signal, not just the scheduling logic on top of the signal.

Karpenter cannot fix inflated requests. It can consolidate pods efficiently onto fewer nodes, but if every pod declares 10x more CPU than it uses, even perfect consolidation leaves most compute idle. The ceiling is the request inflation, not the bin-packing algorithm.

What an autonomous cluster actually does

An autonomous cluster continuously adjusts both layers of the stack based on actual consumption, not declared state. It closes the feedback loop that manual request-setting leaves open. The Cast AI implementation works across three coordinated layers:

Layer 1: Continuous pod rightsizing

OOM kills dropped from 40-50 per 30-minute monitoring window to near zero while provisioned CPUs fell roughly 50%. That result is counterintuitive: reliability improved as resource allocation shrank. It is what happens when requests reflect actual consumption rather than defensive estimates. Smaller requests mean better bin-packing and lower cost. Accurate memory limits mean fewer OOM kills, not more. Rightsizing is not a cost-versus-reliability trade-off; it resolves both problems at once.

Cast AI”s Workload Autoscaler produces this outcome by watching real CPU and memory consumption over time and updating resource requests continuously. It applies changes during natural pod lifecycle events to minimize disruption, not as disruptive mid-run restarts.

Important: when HPA targets CPU utilization as a percentage of the declared request, Cast AI Workload Autoscaler functions as a managed VPA. It can conflict with HPA if both target the same metric. Cast AI addresses this by defaulting to recommendation-only mode when HPA is detected, and raising requests only within the safe zone where HPA target utilization remains achievable. Enable Workload Autoscaler in recommendation-only mode first (managementStrategy: ScaleDown instead of StartupAndOnPolicyUpdate) and validate HPA behavior before enabling full management.

Start with the observation-mode policy below. It applies scale-down corrections only and is safe to run alongside HPA:

# Recommendation-only / scale-down only mode — safe starting point with HPA

apiVersion: cast.ai/v1alpha2

kind: WorkloadScalingPolicy

metadata:

  name: rightsizing-observation

  namespace: default

spec:

  cpuConfig:

    managementStrategy: “ScaleDown”

  memoryConfig:

    managementStrategy: “ScaleDown”

Once you have validated HPA behavior over 7 days, promote to full management with the policy below:

A sample full-management Workload Autoscaler policy:

apiVersion: cast.ai/v1alpha2

kind: WorkloadScalingPolicy

metadata:

  name: production-rightsizing

  namespace: default

spec:

  # applyTypeSelector applies globally to both cpuConfig and memoryConfig

  applyTypeSelector:

    affectedPodState: “Any”

  cpuConfig:

    managementStrategy: “StartupAndOnPolicyUpdate”

    requestIncreaseThreshold: “10%”

    requestDecreaseThreshold: “25%”

    minHeadroom: “5%”

  memoryConfig:

    managementStrategy: “StartupAndOnPolicyUpdate”

    limit: “MaxRecommendation”

    requestIncreaseThreshold: “10%”

    requestDecreaseThreshold: “25%”

This policy applies updated requests at pod startup and after policy changes. Running pods are not disrupted mid-cycle. The system tracks consumption continuously and applies corrections at the next natural restart opportunity.

Layer 2: Coordinated bin-packing (Evictor)

The Cast AI Evictor runs as a continuous bin-packing daemon on top of Karpenter. It does not replace Karpenter; it operates as an overlay. Evictor identifies node fragmentation that Karpenter”s consolidation logic misses, then safely evicts pods to trigger re-scheduling onto fewer, better-utilized nodes.

The benchmark shows this layer adds 9.1% savings on top of Karpenter native. That number understates the value on clusters with high fragmentation from mixed pod sizes and varied scheduling constraints. Evictor respects Pod Disruption Budgets (PDBs) throughout.

One caveat: StatefulSets with EBS volumes are AZ-bound. EBS PersistentVolumes use single-node attachment. The Evictor respects PodDisruptionBudgets but will attempt to drain StatefulSet pods. Ensure PDBs are in place and test Evictor behavior on StatefulSets in staging before enabling on production.

A Helm values override shows the key configuration options:

castai-evictor:

  enabled: true

  aggressiveness: “medium”  # low | medium | high

  disruptionBudgetRespect: true

  nodeGracePeriodMinutes: 10

  scopeSelector:

    matchLabels:

      cast.ai/spot: “false”  # Limit to on-demand nodes; Spot handled separately

The Evictor deploys as a separate Helm chart. To install:

helm repo add castai-helm https://charts.cast.ai

helm install castai-evictor castai-helm/castai-evictor \

  –namespace castai-agent \

  -f evictor-values.yaml

Layer 3: Predictive Spot automation

Spot instances deliver the largest per-node cost reduction available in any cloud. The barrier is interruption risk. Cast AI”s Spot automation uses historical interruption data and real-time market signals to select instance types with low interruption probability, with automatic fallback to on-demand when Spot capacity disappears.

This layer compounds with the first two. Smaller, accurately-sized pods are easier to reschedule on interruption. Better bin-packing means fewer nodes, and fewer nodes means fewer interruption events to manage. The three layers reinforce each other in a way that no single-layer optimization can replicate.

Tool capability comparison

Here is a direct comparison of where each tool operates and what it cannot do:

Capability HPA VPA Karpenter Cluster Autoscaler Cast AI
Scale replica count Yes No No No Yes
Adjust resource requests No Yes (with restart) No No Yes (continuous)
Provision new nodes No No Yes (45-60s) Yes (3-4 min) Yes (via Karpenter)
Consolidate fragmented nodes No No Partial Partial Yes (Evictor)
See actual pod consumption Partial (vs request) Yes No No Yes
Automate Spot selection No No Partial (NodePool) No Yes (predictive)
Run alongside HPA N/A Limited Yes Yes Yes
Zero-disruption rightsizing No No No No Yes (lifecycle-aware)

How to close the gap: a practical starting point

Start with visibility. You cannot fix what you cannot measure. Run the following to see the gap between requested and actual CPU on your current nodes:

# Compare requested vs actual CPU per node

kubectl top nodes

kubectl describe nodes | grep -A5 “Allocated resources”

 

# Check per-pod request vs actual usage

kubectl top pods –all-namespaces –sort-by=cpu

If you see nodes at 60-70% “requested” capacity but 8-15% actual utilization, you have the gap. The next step depends on your current setup and constraints:

  • If you already use Karpenter: Add Cast AI Evictor as an overlay. No changes to your existing NodePool configuration. The Evictor deploys as a separate Helm chart and begins identifying consolidation opportunities immediately.
  • If you use Cluster Autoscaler: Evaluate migrating to Karpenter first, then add Evictor. The provisioning speed difference (45-60s vs 3-4 min) alone justifies the migration for most production workloads.
  • If OOM kills are a recurring problem: Enable Cast AI Workload Autoscaler in recommendation-only mode first. Review the recommendations for 7 days before applying them. This gives your team confidence in the system before any request changes take effect.
  • If Spot is off the table due to interruption risk: Start with on-demand rightsizing and bin-packing. The benchmark shows 15.8% savings with Evictor and CRB alone, with no Spot usage required.

The 43% benchmark result came from a single EKS cluster over 7 days, with rightsizing disabled. Your actual savings depend on your cluster topology, workload mix, and current request inflation. Most teams see measurable results within the first week of enabling rightsizing on non-critical workloads.

 

Connect your cluster for free to Cast AI and see your current waste profile before committing to any configuration changes.

 

Note: Cast AI modifies resource requests in-cluster via webhook, not through your Git manifests. If you use GitOps (ArgoCD, Flux), configure your sync policy to ignore the resources.requests field on namespaces where Cast AI manages requests, or enable drift detection override. Cast AI”s documentation includes an ArgoCD integration guide.

Expert review and source notes

This article was reviewed for technical accuracy against two primary sources: the Cast AI 2026 State of Kubernetes Cost and Usage Report (data from tens of thousands of production clusters) and the Cast AI 7-day adversarial EKS benchmark, published by Cast AI engineering. Benchmark methodology: identical workload and traffic pattern, four Karpenter configurations tested sequentially over 7 days. Cast AI engineering disabled workload rightsizing in the benchmark to isolate the consolidation effect. Actual savings with all layers enabled exceed the 43% figure cited. The 8% average CPU utilization figure represents fleet-wide averages and will vary by workload type and cluster age.

Frequently asked questions

Does Karpenter fix overprovisioning?

No. Karpenter provisions and consolidates nodes based on declared resource requests. If those requests are inflated, and fleet-wide data shows CPU overprovisioning at 69%, Karpenter schedules and bills against the inflated numbers. Karpenter cannot see actual pod consumption. It optimizes bin-packing at the node level, but the underlying signal remains wrong until you correct it with a separate tool like VPA or Cast AI Workload Autoscaler.

 

What is an autonomous cluster?

An autonomous cluster continuously adjusts both pod resource requests and node capacity based on actual consumption data, without manual tuning cycles or incident-driven changes. Cast AI implements this across three coordinated layers: continuous pod rightsizing via the Workload Autoscaler, coordinated node consolidation via the Evictor, and predictive Spot automation. The result is a cluster that self-corrects request drift over time rather than accumulating it.

What is the gap between Karpenter and Cast AI?

Karpenter handles node provisioning and consolidation. It does not adjust resource requests, and it cannot see actual pod consumption. Cast AI adds three capabilities on top: continuous request rightsizing via the Workload Autoscaler, deeper consolidation via the Evictor, and predictive Spot automation. In a 7-day adversarial EKS benchmark, full Cast AI optimization delivered 43% savings versus Karpenter native. That benchmark disabled rightsizing to isolate consolidation, so the real gap is larger.

 

Why does CPU utilization average 8%?

The Cast AI 2026 report, covering tens of thousands of clusters, found fleet-wide average CPU utilization at 8%. This is the compound effect of inflated resource requests, no automated feedback loop to correct them, and autoscalers that provision against declared state rather than actual consumption. CPU overprovisioning reached 69% in 2026, up from 40% the prior year. Without a continuous rightsizing mechanism, clusters accumulate request drift with no self-correction.

 

How does HPA interact with resource requests?

HPA does not modify resource requests. It scales replica count based on a metric, usually CPU or memory utilization as a percentage of the declared request. If a pod declares 500m CPU and uses 40m, HPA sees 8% utilization and holds replica count. The 500m request per pod stays in place. Every downstream decision, including Karpenter node selection and CAS scale-down logic, operates on that uncorrected number. HPA and VPA address different problems and create recommendation conflicts when run on the same metric for the same workload.

 

Can I use Cast AI alongside my existing Karpenter setup?

Yes. Cast AI Evictor runs as an overlay on Karpenter without replacing your existing NodePool configuration. The Workload Autoscaler integrates independently. You can enable each component separately and observe the effect before expanding coverage. Most teams start with Evictor in observation mode, review consolidation recommendations for a week, then enable rightsizing for non-critical workloads first.













Ergänzungen und Infos bitte an die Redaktion per eMail an de-info[at]it-boltwise.de. Da wir bei KI-erzeugten News und Inhalten selten auftretende KI-Halluzinationen nicht ausschließen können, bitten wir Sie bei Falschangaben und Fehlinformationen uns via eMail zu kontaktieren und zu informieren. Bitte vergessen Sie nicht in der eMail die Artikel-Headline zu nennen: "Autoscaling is not optimization: the gap between HPA, Karpenter, and an autonomous cluster".
Alle Märkte in Echtzeit verfolgen - 30 Tage kostenlos testen!

Du hast einen wertvollen Beitrag oder Kommentar zum Artikel "Autoscaling is not optimization: the gap between HPA, Karpenter, and an autonomous cluster" für unsere Leser?

Schreibe einen Kommentar

Deine E-Mail-Adresse wird nicht veröffentlicht. Erforderliche Felder sind mit * markiert

  • Die aktuellen intelligenten Ringe, intelligenten Brillen, intelligenten Uhren oder KI-Smartphones auf Amazon entdecken! (Sponsored)


  • Es werden alle Kommentare moderiert!

    Für eine offene Diskussion behalten wir uns vor, jeden Kommentar zu löschen, der nicht direkt auf das Thema abzielt oder nur den Zweck hat, Leser oder Autoren herabzuwürdigen.

    Wir möchten, dass respektvoll miteinander kommuniziert wird, so als ob die Diskussion mit real anwesenden Personen geführt wird. Dies machen wir für den Großteil unserer Leser, der sachlich und konstruktiv über ein Thema sprechen möchte.

    Du willst nichts verpassen?

    Du möchtest über ähnliche News und Beiträge wie "Autoscaling is not optimization: the gap between HPA, Karpenter, and an autonomous cluster" informiert werden? Neben der E-Mail-Benachrichtigung habt ihr auch die Möglichkeit, den Feed dieses Beitrags zu abonnieren. Wer natürlich alles lesen möchte, der sollte den RSS-Hauptfeed oder IT BOLTWISE® bei Google News wie auch bei Bing News abonnieren.
    Nutze die Google-Suchmaschine für eine weitere Themenrecherche: »Autoscaling is not optimization: the gap between HPA, Karpenter, and an autonomous cluster« bei Google Deutschland suchen, bei Bing oder Google News!


    539 Leser gerade online auf IT BOLTWISE
    KI-Jobs