Agent-Boundary Benchmarks
Product Benchmarks
Agent-Boundary Benchmarks your Chief Information Security Officer can actually reproduce.

Compare Blekline to Industry Leaders
Run DAte //
Date
B1 - Secrets in prompt context
Can raw secrets reach the model context?
Blekline
Pass
Ungoverned
FAIL
Lakera Guard
Pass
Kong AI GW
Partial
B2 - Destructive tools/calls
Are tools/calls with destructive tool args blocked at enforce hop?
Blekline
Pass
Ungoverned
FAIL
Lakera Guard
N/A
Kong AI GW
Partial
B3 - Session lineage after injection
After injection, are destructive tools blocked on contaminated session?
Blekline
Pass
Ungoverned
N/A
Lakera Guard
N/A
Kong AI GW
N/A
B4 - Enforce latency overhead
What is p50/p95/p99 ms on canonical enorce path?
Blekline
0.08 ms
Ungoverned
N/A
Lakera Guard
63 ms
Kong AI GW
306 ms
B5 - Credentials in tool args
Are credentials masked or blocked before tool execution?
Blekline
Pass
Ungoverned
Fail
Lakera Guard
N/A
Kong AI GW
N/A
B6 - Agent pod egress bypass
Can agent pods reach external HTTP without mandatory hop?
Blekline
Pass
Ungoverned
Fail
Lakera Guard
N/A
Kong AI GW
N/A
B7 - Time to first Govern
Minutes from zero to first enforced interaction?
Blekline
Pass
Ungoverned
N/A
Lakera Guard
Partial
Kong AI GW
Partial
B8 - Audit artifact quality
Structured allow/mask/block metadata for SIEM?
Blekline
Pass
Ungoverned
Fail
Lakera Guard
Partial
Kong AI GW
Partial
Frequently Asked Questions
FAQ's
Scoring & methodology
01
What does Pass vs Partial vs Fail mean?
Pass — blocked or masked before model/tool execution. Partial — detected or HTTP/content-only, not structural MCP enforce at the agent hop. Fail — payload reached execution unchanged.
02
Why are some matrix cells N/A?
N/A = outside that product’s layer, not a hidden fail. Lakera classifies prompt content (B1), not MCP tools/call (B2/B3). Kong enforces routes/plugins — no session lineage (B3). OneCLI is doc-verified on egress until local proxy is available.
03
How is audit quality scored in B8?
0–3 scale: 3 = allow/mask/block + findings + requestId; 2 = decision + timestamp; 1 = log/boolean; 0 = no metadata. Blekline scores 3; partial vendors typically 2.
Stack & Competitors
04
Why are Lakera and Kong Partial on some rows?
Lakera flags prompt content — Pass on B1-style injection, N/A on tool enforce. Kong applies route plugins — Partial when the gateway hop exists but tool-arg mask or lineage is out of scope. Partial = detects or routes, not structural enforce at the agent boundary.
05
How is this different from an LLM firewall or prompt filter?
Firewalls inspect text. B1–B8 score the full agent path — mask, destructive tools/call, lineage, egress (B6), time-to-govern (B7), audit metadata (B8). Complementary, not the same layer.
06
Does this replace Kong or Okta?
No. Kong = API routes. Okta/Veza = people and access. Blekline = what agents do at runtime. Regulated teams typically need all three. These benchmarks show where each layer stops.
Reproduction & Diligence
07
Can we reproduce these results in our environment?
Yes, gitignored env.benchmark, pnpm benchmark:run, dated JSON with git SHA. CI runs --quick on every build.
08
Will enforcement slow down our agents?
B4 charts p99. Blekline local enforce targets <10 ms; cloud mask adds RTT. Lakera/Kong include API/gateway hops — compare at the layer each product owns.
09
Who sees our data when we run a reproduction?
Your keys only — never commit secrets. Cloud mask over TLS; audit defaults to metadata only. Sidecar = in-VPC enforce. EEA-first SaaS; US-primary on enterprise addendum.
