Library / Learning / Multi-Armed Bandit

Multi-Armed Bandit

Balance pulling the best-known lever against testing new ones. Regret is the score.

λ mean ≈ variance
counts per window · recreated schematic
INTERACTIVE LAB

Tolerance sandbox (original recreation)

happy 76% • avg. similar neighbors 49% • red ring = wants to move

LEARNING

When to use / misuse

Use → Trials, dosing, ad creative. Run small experiments, then exploit.

Difficulty → Core

REDCAPE → Act · Explore

agent: bandit-checker
family: Learning
combine_with: [two other families]
rule: state failure mode before acting

Free sources: Book final chapter (opioid case) — original summary