---
title: "Chapter 3: High-Density Kubernetes & Karpenter Autoscaling — Cloud Economics | TinyCTO"
description: "Just-in-Time (JIT) node provisioning, Graviton ARM64 adoption, Spot consolidation, and bin-packing efficiency."
image: "https://tinycto.tv/assets/cloud-economics/cloud_economics_manuals_og.jpg"
canonicalUrl: "https://tinycto.tv/cloud-economics/manuals/03-compute-karpenter-autoscaling"
locale: "en"
---

# Chapter 3: High-Density Kubernetes & Karpenter Autoscaling

## 1. Executive Summary & Problem Statement
Legacy Kubernetes Cluster Autoscaler (CAS) scales worker nodes based on rigid Auto Scaling Groups (ASGs). When pods become unschedulable, CAS provisions predefined node sizes that frequently lead to extreme resource fragmentation: a single 500m CPU pod can trigger an m5.4xlarge node spinning up, resulting in $85\%$ idle compute capacity.

**Karpenter** introduces Just-in-Time (JIT) node provisioning. Instead of matching predefined ASG templates, Karpenter directly negotiates with cloud APIs (AWS Fleet API) to launch the optimal mix of Graviton (ARM64) instances, Spot instances, and variable instance types tailored to pending pod requirements.

---

## 2. Karpenter NodePool Specification
A production Karpenter NodePool configuration prioritizing ARM64 Graviton instances, Spot preemptible pools, and automated node consolidation:

```yaml
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: general-compute-finops
spec:
  template:
    spec:
      requirements:
        - key: kubernetes.io/arch
          operator: In
          values: ["arm64"] # Graviton3 / Graviton4 (20% better price-performance)
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"] # Prefers Spot; falls back to on-demand
        - key: node.kubernetes.io/instance-type
          operator: In
          values: ["c7g.xlarge", "c7g.2xlarge", "m7g.xlarge", "m7g.2xlarge", "r7g.xlarge"]
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: default-al2023
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m # Consolidate pods to smaller nodes within 60s of underutilization
    expireAfter: 720h    # 30-day node rotation to prevent OS drift
```

---

## 3. Bin-Packing Efficiency & Consolidation Economics
Karpenter continuously watches cluster utilization. When traffic subsides:
1. It identifies underutilized nodes where pods can fit onto existing nodes.
2. It provisions a smaller instance (or drains onto existing nodes).
3. It cordons and terminates the underutilized node, releasing the EC2 compute charges within 60 seconds.

$$\text{Bin Packing Efficiency} = \frac{\sum \text{Pod Requests (CPU / Memory)}}{\sum \text{Allocatable Node Capacity (CPU / Memory)}} \times 100$$

A healthy Karpenter cluster maintains $> 82\%$ average CPU and memory utilization, compared to $35\% - 50\%$ in traditional static clusters.

---

## 4. Best Practices & Invariants
- **Set Realistic Resource Requests:** Pods with bloated CPU requests prevent efficient bin-packing. Enforce Vertical Pod Autoscaler (VPA) in recommendation mode to rightsize requests based on p95 actual usage.
- **Topology Spread Constraints:** Use `topologySpreadConstraints` with `maxSkew: 1` across availability zones to prevent Karpenter from scheduling all Spot instances into a single AZ prone to Spot interruptions.

```json
{
  "@context": "https://schema.org",
  "@type": "TechArticle",
  "headline": "Chapter 3: High-Density Kubernetes & Karpenter Autoscaling",
  "description": "Just-in-Time (JIT) node provisioning, Graviton ARM64 adoption, Spot consolidation, and bin-packing efficiency.",
  "url": "https://tinycto.tv/cloud-economics/manuals/03-compute-karpenter-autoscaling",
  "inLanguage": "en-US",
  "author": {
    "@type": "Organization",
    "name": "TinyCTO.tv",
    "url": "https://tinycto.tv"
  },
  "publisher": {
    "@type": "Organization",
    "name": "TinyCTO.tv",
    "url": "https://tinycto.tv"
  }
}
```
