- 01
Test files, not real invoices
The invoices come from a test suite. They are built to exercise rules and say nothing about how often a detail appears only in the text of real invoices. Some test cases are variants of the same invoice with the same free text, so a finding in several rows can go back to the same wording.
- 02
English questions about German text
The questions to Jev are in English, the invoice texts in German. We did not measure German questions against them.
- 03
One run
Each verdict comes from a single run. In an earlier measurement on the same test invoices, with different questions, a second run differed by at most 0.09.
- 04
Cross-check only for conspicuous cells
At the default thresholds a second language model read every cell in the classes "text only" and "uncertain" against the wording of the invoice. A third, independent check then overturned three cells, which the detail panel marks as refuted. No person has reviewed this cross-check yet. A month in which goods were delivered counts as a delivery time, not as a service period. Consistent cells were not checked. If you move the thresholds, the detail panel shows which cells are unchecked.
- 05
Twelve line items at most
For invoices with more than twelve line items, Jev saw the texts of the first twelve only. Anything further down was not part of the request.
- 06
Replay, not a live call
The page plays back stored results. Your browser sends no request to a model.
- 07
Origin and licence
The test invoices come from the XRechnung test suite by KoSIT, published under the Apache-2.0 licence. The price per million input tokens is listed in the TypeSafe model overview.
XRechnung test suite on GitHub ↗TypeSafe model overview ↗