660 accounting cases.
660 correct.
Before AI touches your books, the maths underneath has to be right. We test Workpaper's bank, GST and TDS engines on randomly generated Indian company data, against answers worked out by hand from accounting rules and the tax law.
100%of 660 cases correct
0confidently wrong answers
100%of bank matches correct
100%of true bank matches found
By area
| Area | Scenarios | Cases | Correct |
|---|---|---|---|
| Bank and card reconciliation | 11 | 220 | 100% |
| GST input credit vs GSTR-2B | 10 | 200 | 100% |
| TDS: section, rate and threshold | 12 | 240 | 100% |
By difficulty
| Kind of case | What it tests | Cases | Correct |
|---|---|---|---|
| normal | Clean, everyday data | 120 | 100% |
| messy | Real-world noise: punctuation, timing gaps, duplicates | 200 | 100% |
| ambiguous | Two plausible answers: it must refuse to guess | 60 | 100% |
| contradictory | Sources disagree: it must flag, not force a match | 40 | 100% |
| dangerous | Lookalikes that would cause a wrong entry or overclaim | 140 | 100% |
| human | Needs a person's judgment: it must stop and ask | 20 | 100% |
| law | Recent changes in Indian tax law | 80 | 100% |
Every scenario
Each scenario runs on 20 different randomly generated datasets: new amounts, GSTINs, PANs, UTRs and dates each time.
Bank and card reconciliation
- Same UTR, amount and daynormal20/20
- One UTR written differently in Tally and the bank exportmessy20/20
- No reference; bank clears two days after the booksmessy20/20
- Two equally likely bank lines for one entryambiguous20/20
- Same amount 27 days apart: a coincidence, not a matchdangerous20/20
- Same amount, different UTRsdangerous20/20
- Bank export repeats a linemessy20/20
- Customer paid net of 10% TDScontradictory20/20
- Digits swapped: ₹54,321 booked, ₹45,321 in the bankdangerous20/20
- A bounced receipt shows as a debitdangerous20/20
- Bank charge with no book entrymessy20/20
GST input credit vs GSTR-2B
- Invoice matches 2B exactlynormal20/20
- INV-1234 in the books, INV/1234 in 2Bmessy20/20
- Tax differs by under ₹1normal20/20
- Supplier hasn't reported the invoicemessy20/20
- 2B shows ₹500 less taxcontradictory20/20
- Valid-looking but wrong GSTINdangerous20/20
- Same invoice number, different supplierambiguous20/20
- One bill booked twicedangerous20/20
- One malformed GSTIN in a batchmessy20/20
- Bill in 2B that was never bookedmessy20/20
TDS: section, rate and threshold
- Professional fees of ₹1,00,000normal20/20
- Contractor payment to a companynormal20/20
- Single contract below ₹30,000normal20/20
- Contractor crosses ₹1,00,000 in the yearmessy20/20
- Payee without a PANdangerous20/20
- Individual rate chosen for a company PANambiguous20/20
- Non-filer payee in FY 2024-25messy20/20
- Non-filer payee after section 206AB was omittedlaw20/20
- Payment to a non-resident without a reviewed ratehuman20/20
- Professional fees of ₹40,000 under the new ₹50,000 thresholdlaw20/20
- Rent of ₹60,000 against the ₹50,000-a-month thresholdlaw20/20
- Commission after the cut to 2%law20/20
Guardrails on AI-drafted entries
Before any journal the AI drafts is saved, Workpaper replays it against your bank statement and your books. Our test suite checks that it:
- blocks an entry that would inflate the bank balance with no matching statement line;
- blocks a correction that has already been made once;
- blocks any entry that touches share capital or reserves;
- allows a genuine correction, such as booking a ₹590 bank charge, even when other differences exist.
Workpaper does not post to your ledger or file returns. People approve, post and file.
How it works, and its limits
- Hand-derived answers. Expected results come from accounting rules and the Income-tax Act, written separately from the engines. They are never copied from engine output.
- What counts as wrong. A wrong match, an input credit overclaim or the wrong tax amount. Refusing to guess on an ambiguous case counts as correct.
- Law as of FY 2026-27, including Finance Act 2025.
- Engines, not the AI. This measures the deterministic calculations the AI relies on. We test the AI agent separately on full synthetic company-months, and will publish those results once they're stable.
- Synthetic data. No customer data is used.