jev-pl-benchmark

Typed-decision systems on Polish and English decisions. Accuracy with both option orders averaged (Polish decisions scored with the corrected notice-period labels of the technical report, v1.0.1); speed over HTTP on one H100 80GB, each system served with its own official server. Test items are private.

Columns

polish_decisions
7,081 held-out Polish decisions (deadlines, amounts, rule families, statutes, broad decisions) from unseen templates, statutes and domains
polish_general
3,079 items: Polish knowledge and exam questions, reading comprehension and entailment, synthetic decisions, English control
english_decisions
1,479 held-out English decisions
public_english_bench
231-item public English decision benchmark, official harness
latency_p50_ms_h100
median latency of sequential requests over HTTP (ms)
requests_per_s_h100
requests per second with 32 concurrent clients; one request per item in the original option order
polish_decisions_one_request
accuracy of exactly the timed single request on the Polish decisions. The accuracy columns average two requests (original and reversed option order) for every system except basal, which averages both orders inside one request

Empty speed cells: commercial API (network and hardware unknown) or no official HTTP server.

basal-1.0 rows are measured with the released basal engine v1.0.1 (basal-serve --mode fast, bf16, both option orders), with the same evaluation and load-test tools as every other system. Latency: median of 200 sequential requests; throughput: 32 concurrent clients. These are measured on the benchmark sets; on the 500-item speed sample of HARDWARE.md the same server gives 14.1 ms and 63 requests/s (basal-1.0-4.5B). More speed results (B300, RTX PRO 6000, RTX 5090, RTX 4090, DGX Spark; FP8, NVFP4, early exit): HARDWARE.md.