---
title: "Chapter 02: Raft & Paxos Consensus: Leader Elections, Term Epochs & State Machine Replication | TinyCTO Distributed Systems Canon"
description: "Internal mechanics of formally verified consensus protocols: randomized election timers, split-vote mitigation, Pre-Vote probing, log matching invariants, and lease-based read optimizations."
image: "https://tinycto.tv/assets/distributed-systems/distributed_systems_manuals_og.jpg"
canonicalUrl: "https://tinycto.tv/distributed-systems/manuals/02-raft-paxos-consensus-internals"
locale: "en"
---

# Chapter 02: Raft & Paxos Consensus: Leader Elections, Term Epochs & State Machine Replication

> **Canonical Distributed Systems Engineering Field Manual**
> **Theorem Citation**: Ongaro & Ousterhout (Raft 2014) / Leslie Lamport (Paxos 1998) | **Read Time**: 22 min read | **Maturity Target**: SCALED

Internal mechanics of formally verified consensus protocols: randomized election timers, split-vote mitigation, Pre-Vote probing, log matching invariants, and lease-based read optimizations.

# Raft & Paxos Consensus: Leader Elections, Term Epochs & State Machine Replication

## Executive Summary
Replicated state machines are the foundation of strongly consistent distributed infrastructure. By ensuring that a cluster of independent nodes executes the exact same sequence of deterministic state transitions, the cluster acts as a single, fault-tolerant entity. Ongaro and Ousterhout's Raft protocol decomposed consensus into understandable subproblems: leader election, log replication, and safety invariants.

## 1. The Quorum Formula
To tolerate $F$ node crashes, a Raft cluster must contain at least $2F + 1$ nodes. The minimum quorum size $Q$ required to elect a leader or commit a log entry is:
$$Q = \left\lfloor \frac{N}{2} \right\rfloor + 1$$
For a 5-node cluster ($N=5$), $F=2$, and $Q=3$. Any two quorums of size $Q$ must overlap in at least one node, guaranteeing that the winning candidate has seen every previously committed entry.

```mermaid
sequenceDiagram
    autonumber
    participant C as Client
    participant L as Raft Leader (Term 4)
    participant F1 as Follower 1
    participant F2 as Follower 2

    C->>L: 1. Write Command ("set key=val")
    Note over L: Append to Local Log<br/>Term: 4, Index: 104
    L->>F1: 2. AppendEntries RPC (Term 4, Index 104)
    L->>F2: 2. AppendEntries RPC (Term 4, Index 104)
    F1-->>L: 3. AppendEntries Success ACK
    Note over L: Quorum Reached (2/3 ACKs)<br/>Commit Index -> 104
    L->>C: 4. Success Response
    L->>F2: 5. Heartbeat (CommitIndex=104)
```

## 2. Raft Leader Election & Term Epochs
- **Randomized Election Timeouts (150ms - 300ms):** Prevents split-vote livelocks where multiple candidates start elections simultaneously.
- **Pre-Vote Extension:** Candidates first send low-overhead Pre-Vote probes without incrementing their term. If they cannot reach a majority, they do not trigger a disruptive cluster-wide re-election.
- **Leader Completeness Property:** A voter denies its vote if the candidate's log is less up-to-date than its own log ($	ext{Term}_{candidate} < 	ext{Term}_{voter}$ or $	ext{Index}_{candidate} < 	ext{Index}_{voter}$).

## 3. Fencing Tokens: Eliminating Zombie Leaders
When a network partition temporarily isolates a leader, it may experience a "zombie" state where it still believes it is the leader while the rest of the cluster has elected a new leader with a higher term epoch.
Fencing tokens (monotonically increasing integer term numbers) are passed with every downstream RPC:
```
Client Request -> Leader (Term 5, FencingToken: 5) -> Storage
New Leader Elected (Term 6, FencingToken: 6) -> Storage rejects Term 5 RPCs with FENCING_TOKEN_STALE
```


### Core Concepts & Consistency Models

- `Raft State Machine`
- `Multi-Paxos`
- `Term Epochs`
- `Pre-Vote Protocol`
- `Fencing Tokens`

### Canon Surfaces & Navigation

- **Manuals Library**: https://tinycto.tv/distributed-systems/manuals
- **18 Reference Architectures**: https://tinycto.tv/distributed-systems/architectures
- **Topology & Sizer Wizard**: https://tinycto.tv/distributed-systems/wizard
- **Technology & Consensus Matrix**: https://tinycto.tv/distributed-systems/matrix

```json
{
  "@context": "https://schema.org",
  "@type": "TechArticle",
  "headline": "Raft & Paxos Consensus: Leader Elections, Term Epochs & State Machine Replication",
  "description": "Internal mechanics of formally verified consensus protocols: randomized election timers, split-vote mitigation, Pre-Vote probing, log matching invariants, and lease-based read optimizations.",
  "inLanguage": "en",
  "educationalLevel": "Advanced",
  "proficiencyLevel": "Expert",
  "url": "https://tinycto.tv/distributed-systems/manuals/02-raft-paxos-consensus-internals"
}
```
