SHARED STATE
Agent recovery trace
Attempt three changed the same files. The same two tests still fail. A reviewer says the working hypothesis no longer explains the evidence.
- Attempts3
- Repeated failureYes
- BudgetMedium
SHARED-STATE DECISION MODEL
Give Blitz one view of the situation. Ask a panel of focused questions. Get explicit distributions that software can compose into its next action—without generating a paragraph.
01 / SHARED-STATE FANOUT
Blitz acts like a semantic sensor panel. Each question reads the same evidence independently; ordinary code combines the resulting distributions.
SHARED STATE
Attempt three changed the same files. The same two tests still fail. A reviewer says the working hypothesis no longer explains the evidence.
YES91%
NO86%
YES78%
NO64%
CHANGE STRATEGY82%
CODE COMPOSES
Stop repeating the failed approach and form a new hypothesis.
Illustrative probabilities. The live playground below exposes the atomic decision primitive; shared-state batched serving is the next measured milestone.
02 / INSPECT THE PRIMITIVE
A workflow is built from small, focused judgments. Start with a real checkpoint example, then change the evidence, question, or options.
Your decision appears here
Selected decision
Decision receipt
03 / MEASURED, NOT PROMISED
Fresh confirmation results and a warm A100 throughput measurement from the current research checkpoint.
Synthetic held-out evaluation; not a claim of real-world accuracy or calibration.
04 / THE MODEL'S JOB
Blitz is a compact 0.6B model adapted to score bounded choices in one forward pass per atomic question. It handles the semantic judgment between observed evidence and a software system's next action.
Present one concise state that every decision can inspect.
Evaluate independent routing, evidence-availability, risk, and recovery signals.
Use explicit distributions to continue, retry, review, reroute, or stop.
WHERE BLITZ FITS
Detect repeated failures and choose whether to retry, review, or change strategy.
Judge severity, route ownership, and decide when ambiguous evidence needs escalation.
Score retrieval relevance, context coverage, semantic equivalence, and experiment outcomes at scale.