Chapter 01 / Security automation
Arbiter AI.
A local-first approach to security alert triage.
Recorded evaluation · August 22, 2026
- Cases across two evaluated suites
- 121
- Recall · realistic suite
- 100%
- Precision · realistic suite
- 72%
- Python versions in CI configuration
- 4
In the recorded realistic-suite comparison, Qwen caught all labeled attacks while achieving 83% overall accuracy. The evaluation makes the tradeoff visible: preserving recall still leaves false positives for review.
Measurement scope & source
Qwen3.5:4b with guardrails enabled; 101 realistic cases plus 20 red-team cases, three runs per suite. Recall, precision and accuracy above refer only to the full realistic suite, including rule decisions—not model-only performance or production detection. Local Ollama inference. CI config covers Python 3.10–3.13; its mock evaluation is non-blocking. This historical benchmark was not rerun for the portfolio.
The idea
Rules for the clear-cut. Local AI for the ambiguous. A traceable decision for every alert.
Python / Ollama / SQLite / GitHub Actions
Animated architecture
Inside the triage engine
Branch labels show alternative routes. Playback highlights architectural stages, not a live security event.
01 / Receive
02 / Enforce the boundary
03 / Review ambiguous events
04 / Check the proposed decision
05 / Preserve the record
Control flow / Decisions & data
How an event becomes a verdict
Follow the branches: immediate escalation, deterministic decisions, or a contextual model review. All completed decisions converge on the audit trail.
Scroll across the diagram to follow each branch
Reading the flow
- An unexcepted guardrail hit escalates before inference. A corroborating administrator fact can route that specific hit to model review; it never grants suppression by itself.
- A clear pre-filter verdict skips the model. Ambiguous events receive signature history, asset criticality, and scope-checked facts.
- Model failure, a suppression confidence below 0.7, or an applicable post-model guardrail veto leads to escalation. The matching-fact exception is retained in the rationale.
- Completed verdicts update memory and append a JSONL record. Shadow mode is enabled by default; this diagram does not imply an autonomous response action.
Why / What / How
The thinking behind the system.
Why a hybrid triage engine?
A queue of alerts is not yet a queue of decisions. An analyst still needs to connect the event to the affected asset, previous observations and the rules that apply. Arbiter explores how to make that first pass repeatable while keeping its reasoning available for review.
Deterministic rules handle cases where the system already has enough information. Local inference is reserved for ambiguity. This avoids asking a model to reconsider every clear rule match and gives the model a narrower job: interpret context inside boundaries enforced by code.
What the system actually produces
The central output is a recorded verdict with its rationale and supporting evidence. SQLite supplies historical context and stores decisions; the audit record preserves how an event was handled. Dashboard views expose that record through a JSON API.
The demonstration uses sample events. Shadow mode and response dry-run defaults let the decision path be inspected before any operational response is considered. A suggested verdict is therefore distinct from an action taken on a host.
How context enters without becoming permission
The engine checks dangerous patterns before inference, then lets the pre-filter resolve clear cases. An ambiguous event receives scoped history, asset criticality and administrator-curated facts before reaching Ollama.
A matching curated fact can send a specific guardrail hit for model review, but does not grant suppression by itself. After inference, the suppression threshold and applicable guardrail veto still matter. Model failure or insufficient confidence routes the event to review instead of silently discarding it.
A concrete route through the design
Consider an event whose meaning depends on whether the activity was expected on that host. With no decisive rule match, the engine gathers context and asks the local model for a verdict. A request to suppress with confidence below 0.7 becomes an escalation; an accepted verdict is recorded with its rationale.
This example explains the implemented control path, not a measured detection result. The repository’s evaluation artifacts help inspect particular failure modes, including an attack described as maintenance; broader collector integration and operational validation remain separate work.
The starting point
Why this project?
Small organizations can generate security alerts without having a dedicated SOC to investigate them. Arbiter explores a self-hosted triage workflow that keeps event analysis local and makes each decision inspectable.
My contribution
The work I brought to it.
I am building Arbiter as a personal security project: exploring how deterministic rules, local inference, environment context, and an audit trail can work together. The central design question is what the model should be allowed to decide—and where code must enforce the boundary.
How it works
From input to output.
Events pass through security guardrails and a deterministic pre-filter. Ambiguous events reach an Ollama-backed local model with asset context and verdict history from SQLite. The engine records a rationale and evidence with its verdict. Inference failures and low-confidence suppression requests escalate for review. Shadow mode and response dry-run are defaults. A JSON API connects the audit store to React dashboard views.
- Sample events
- Guardrails + pre-filter
- Local model, when needed
- Verdict + audit trail
A design decision
Give the model context. Keep the guardrails in code.
Adversarial evaluation exposed a failure mode: an attack framed as routine maintenance could persuade a model to suppress it. Non-suppressible guardrails check dangerous patterns before and after inference. Current code also permits narrowly scoped, administrator-curated facts to route a matching guardrail to model review; that exception is recorded in the rationale. This trades some flexibility and extra review for a more inspectable safety boundary.
The evidence
Look under the surface.
These links point to the reviewed source revision, so the implementation behind this story stays inspectable.
Triage engine
Inspect escalation on model failure, the suppression confidence threshold, and recorded guardrail exceptions.
arbiter/triage.pyGuardrail tests
Review executable cases around dangerous patterns and attempts to explain them away.
tests/test_guardrails.pyModel comparison
An August 2026 experiment separates model-only decisions from guardrail decisions. Results are bounded to those labeled suites and that setup.
evals/2026-08-22-1536/comparison.mdPortable demo guide
Build instructions describe a local sample-event demonstration with bundled runtime and model.
docs/PORTABLE.mdCurrent boundaries
Useful work. Honest limits.
- Pre-release software, not a claim of production readiness or universal detection.
- Dashboard views and a portable builder exist in source. Clean-machine, USB, and broader hardware validation remain open.
- The demonstration replays sample events. Real host collectors and a signed general release remain future work.
- Evaluation artifacts are available for inspection; no benchmark was rerun for this portfolio.
Source reviewed October 3, 2026.
Let’s build something
worth protecting.
“A thoughtful conversation can be the beginning of something worth building.”Connect on LinkedIn