correct decisions (65/65 gold-standard cases)
Measured, not claimed.
Compliance decisions must not depend on chance. That is why Agentic360 makes regulatory decisions not with a language model but with a deterministic policy engine (Open Policy Agent) — and we measure it: across five regulatory domains, with 65 legally annotated test cases including boundary and exemption cases.
latency per decision (p99, single core)
deterministic — same question, same answer
What we measure
The benchmark covers e-invoicing (EN 16931/XRechnung), the EU AI Act, EEG/EnWG, NIS2 and GDPR — plus the affected-party engine that determines which tenants are hit by every regulatory change. Five metrics, all reproducible:
Correctness
Every test case is a realistic profile with a legally expected decision — including boundary cases such as 49 vs. 50 employees (NIS2) or 100.0 vs. 100.1 kWp (EEG). No false positives, no false negatives.
Coverage
Share of the policy logic actually exercised by the test corpus — the rules are not just written, they are verified.
Latency & throughput
Decisions per second on a single CPU core, without GPU and without cloud API. p99 below 0.7 milliseconds.
Determinism
Five identical runs per domain, byte-identical results, SHA-256-verified. The property no purely generative system can offer.
Explainability
Every positive decision names rule and reason — machine-readable, fed directly into the audit trail.
Latency per regulatory domain
Measured with OPA on a single CPU core — the most conservative configuration. In practice: one regulatory change is checked against 10,000 tenant profiles in under 2 seconds.
| Domain | Mean ms | p99 ms | Decisions/s |
|---|---|---|---|
| Affected-party analysis (Affected Engine) | 0.19 | 0.66 | 5,351 |
| EU AI Act (risk classification) | 0.13 | 0.52 | 7,580 |
| E-invoicing (EN 16931 / XRechnung) | 0.14 | 0.55 | 6,999 |
| EEG / EnWG § 14a | 0.13 | 0.42 | 7,839 |
| GDPR | 0.13 | 0.47 | 7,752 |
| NIS2 (scoping) | 0.14 | 0.48 | 7,227 |
Why we don’t leave decisions to the LLM
The language model understands your question and your documents. The verdict comes from the policy engine. This division of labor is measurably better:
| LLM-only (general-purpose chatbots) | Agentic360 (hybrid) | |
|---|---|---|
| Same question → same answer? | not guaranteed | always — proven |
| Compliance check response time | seconds | under 1 millisecond |
| Boundary cases (e.g. 49 vs. 50 employees) | error-prone | 100% correct in test |
| Verifiable for auditors & regulators? | no — prose | yes — rule + reason |
| Does data leave your premises? | usually yes | no — local, CPU |
| Cost per check | API fees | ≈ €0 |
Transparency
All figures come from a reproducible measurement run (July 2026, OPA) against an internally annotated gold standard of 65 cases. “100% correctness” refers to this test corpus — not to every conceivable situation — and does not replace legal advice. We provide the full test suite on request: recalculating is explicitly encouraged.
Recalculating encouraged.
We will show you the benchmark live — with your own scenarios if you like. 30 minutes, no sales loop.
Request a demo