Agent-Boundary Benchmarks

Product Benchmarks

Agent-Boundary Benchmarks your Chief Information Security Officer can actually reproduce.

Compare Blekline to Industry Leaders

Run DAte //

Date

B1 - Secrets in prompt context

Can raw secrets reach the model context?

Blekline

Pass

Ungoverned

FAIL

Lakera Guard

Pass

Kong AI GW

Partial

B2 - Destructive tools/calls

Are tools/calls with destructive tool args blocked at enforce hop?

Blekline

Pass

Ungoverned

FAIL

Lakera Guard

N/A

Kong AI GW

Partial

B3 - Session lineage after injection

After injection, are destructive tools blocked on contaminated session?

Blekline

Pass

Ungoverned

N/A

Lakera Guard

N/A

Kong AI GW

N/A

B4 - Enforce latency overhead

What is p50/p95/p99 ms on canonical enorce path?

Blekline

0.08 ms

Ungoverned

N/A

Lakera Guard

63 ms

Kong AI GW

306 ms

B5 - Credentials in tool args

Are credentials masked or blocked before tool execution?

Blekline

Pass

Ungoverned

Fail

Lakera Guard

N/A

Kong AI GW

N/A

B6 - Agent pod egress bypass

Can agent pods reach external HTTP without mandatory hop?

Blekline

Pass

Ungoverned

Fail

Lakera Guard

N/A

Kong AI GW

N/A

B7 - Time to first Govern

Minutes from zero to first enforced interaction?

Blekline

Pass

Ungoverned

N/A

Lakera Guard

Partial

Kong AI GW

Partial

B8 - Audit artifact quality

Structured allow/mask/block metadata for SIEM?

Blekline

Pass

Ungoverned

Fail

Lakera Guard

Partial

Kong AI GW

Partial

Frequently Asked Questions

FAQ's

Scoring & methodology

01

What does Pass vs Partial vs Fail mean?

Pass — blocked or masked before model/tool execution. Partial — detected or HTTP/content-only, not structural MCP enforce at the agent hop. Fail — payload reached execution unchanged.

02

Why are some matrix cells N/A?

N/A = outside that product’s layer, not a hidden fail. Lakera classifies prompt content (B1), not MCP tools/call (B2/B3). Kong enforces routes/plugins — no session lineage (B3). OneCLI is doc-verified on egress until local proxy is available.

03

How is audit quality scored in B8?

0–3 scale: 3 = allow/mask/block + findings + requestId; 2 = decision + timestamp; 1 = log/boolean; 0 = no metadata. Blekline scores 3; partial vendors typically 2.
Stack & Competitors

04

Why are Lakera and Kong Partial on some rows?

Lakera flags prompt content — Pass on B1-style injection, N/A on tool enforce. Kong applies route plugins — Partial when the gateway hop exists but tool-arg mask or lineage is out of scope. Partial = detects or routes, not structural enforce at the agent boundary.

05

How is this different from an LLM firewall or prompt filter?

Firewalls inspect text. B1–B8 score the full agent path — mask, destructive tools/call, lineage, egress (B6), time-to-govern (B7), audit metadata (B8). Complementary, not the same layer.

06

Does this replace Kong or Okta?

No. Kong = API routes. Okta/Veza = people and access. Blekline = what agents do at runtime. Regulated teams typically need all three. These benchmarks show where each layer stops.
Reproduction & Diligence

07

Can we reproduce these results in our environment?

Yes, gitignored env.benchmark, pnpm benchmark:run, dated JSON with git SHA. CI runs --quick on every build.

08

Will enforcement slow down our agents?

B4 charts p99. Blekline local enforce targets <10 ms; cloud mask adds RTT. Lakera/Kong include API/gateway hops — compare at the layer each product owns.

09

Who sees our data when we run a reproduction?

Your keys only — never commit secrets. Cloud mask over TLS; audit defaults to metadata only. Sidecar = in-VPC enforce. EEA-first SaaS; US-primary on enterprise addendum.

Map one production agent workflow, staging sidecar or MCP proxy, audit export included.