Skip to main content

> FINOPS // CHAPTER 03

Chapter 3: High-Density Kubernetes & Karpenter Autoscaling

Just-in-Time (JIT) node provisioning, Graviton ARM64 adoption, Spot consolidation, and bin-packing efficiency.

Canonical FinOps Manual #03|TinyCTO Cloud Bill Bible

Chapter 3: High-Density Kubernetes & Karpenter Autoscaling

Just-in-Time (JIT) node provisioning, Graviton ARM64 adoption, Spot consolidation, and bin-packing efficiency.

#1. Executive Summary & Problem Statement

Legacy Kubernetes Cluster Autoscaler (CAS) scales worker nodes based on rigid Auto Scaling Groups (ASGs). When pods become unschedulable, CAS provisions predefined node sizes that frequently lead to extreme resource fragmentation: a single 500m CPU pod can trigger an m5.4xlarge node spinning up, resulting in 85%85\% idle compute capacity.

Karpenter introduces Just-in-Time (JIT) node provisioning. Instead of matching predefined ASG templates, Karpenter directly negotiates with cloud APIs (AWS Fleet API) to launch the optimal mix of Graviton (ARM64) instances, Spot instances, and variable instance types tailored to pending pod requirements.


#2. Karpenter NodePool Specification

A production Karpenter NodePool configuration prioritizing ARM64 Graviton instances, Spot preemptible pools, and automated node consolidation:

apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: general-compute-finops
spec:
  template:
    spec:
      requirements:
        - key: kubernetes.io/arch
          operator: In
          values: ["arm64"] # Graviton3 / Graviton4 (20% better price-performance)
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"] # Prefers Spot; falls back to on-demand
        - key: node.kubernetes.io/instance-type
          operator: In
          values: ["c7g.xlarge", "c7g.2xlarge", "m7g.xlarge", "m7g.2xlarge", "r7g.xlarge"]
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: default-al2023
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m # Consolidate pods to smaller nodes within 60s of underutilization
    expireAfter: 720h    # 30-day node rotation to prevent OS drift

#3. Bin-Packing Efficiency & Consolidation Economics

Karpenter continuously watches cluster utilization. When traffic subsides:

  1. It identifies underutilized nodes where pods can fit onto existing nodes.
  2. It provisions a smaller instance (or drains onto existing nodes).
  3. It cordons and terminates the underutilized node, releasing the EC2 compute charges within 60 seconds.
Bin Packing Efficiency=∑Pod Requests (CPU / Memory)∑Allocatable Node Capacity (CPU / Memory)×100\text{Bin Packing Efficiency} = \frac{\sum \text{Pod Requests (CPU / Memory)}}{\sum \text{Allocatable Node Capacity (CPU / Memory)}} \times 100

A healthy Karpenter cluster maintains >82%> 82\% average CPU and memory utilization, compared to 35%−50%35\% - 50\% in traditional static clusters.


#4. Best Practices & Invariants

  • Set Realistic Resource Requests: Pods with bloated CPU requests prevent efficient bin-packing. Enforce Vertical Pod Autoscaler (VPA) in recommendation mode to rightsize requests based on p95 actual usage.
  • Topology Spread Constraints: Use topologySpreadConstraints with maxSkew: 1 across availability zones to prevent Karpenter from scheduling all Spot instances into a single AZ prone to Spot interruptions.
AI Summary — Chapter 03: Chapter 3: High-Density Kubernetes & Karpenter Autoscaling
AEO / GEO / Perplexity Indexable

Just-in-Time (JIT) node provisioning, Graviton ARM64 adoption, Spot consolidation, and bin-packing efficiency.

Chapter ScopeChapter 03 canonical FinOps principles and unit cost guardrails.
Core ConceptsKarpenter NodePool • ARM64 Graviton • Bin-Packing Efficiency • Automated Consolidation
Maturity LevelWALK (Intermediate)
Agent GuardrailEnforce FOCUS 1.0 mandatory tagging schema and automated anomaly gate remediation.