TruvaLI technical report
This report describes how TruvaLI is built: the source archive, how training examples are generated and filtered, how the model is trained and served, how v1.1 scores against its base model, where it fails, and the roadmap to June 2027.
Version 1.1 · October 2026 · Bitrelic Teknoloji A.Ş.
Abstract
General-purpose language models are weak at the parts of compliance work that matter most: they invent article numbers and deadlines, they do not know Turkish regulation in depth, and they run on servers an institution cannot send customer data to. We describe TruvaLI, a vision-language model adapted to AML, fraud and compliance from an archive of about 6,000 sources. We fine-tuned an open-source foundation model with LoRA on synthetic examples generated by self-hosted open-source teacher models and filtered through domain-specific quality gates. Against its base model, v1.1 raises compliance exam accuracy from 88.8% to 92.6%, grounded question answering from 89.3% to 91.6%, and valid structured output from 0% to 89.3%. We report its failures alongside these results and set out the roadmap to June 2027.
1. Introduction
A compliance officer's day is reading: regulation, typology reports, case files, statements, adverse media. Large language models read quickly, but three problems keep them out of the work. They answer confidently when they do not know, which is unacceptable where a wrong deadline is a regulatory breach. Their knowledge of Turkish regulation (MASAK, BDDK, TCMB, SPK, KVKK) is thin. And most of them only run on a provider's servers, where customer data cannot go.
TruvaLI set out to answer three questions:
- Can an open model, adapted with domain data, beat its base model on compliance tasks while staying honest about what it does not know?
- Can the training data be produced without outputs from closed commercial models and without raw customer data?
- Can the result run on a single server that the institution or we control?
2. Design principles
- Grounded: an answer about regulation cites the passage it came from, or says the passage is not there.
- Masked: personal data appears only as
MASKED:tokens; an example containing unmasked data is dropped. - Clean lineage: synthetic data comes from open-source teacher models on our own hardware; closed commercial models are used only for design and evaluation.
- Jurisdiction-aware: a foreign rule is named as foreign, never presented as Turkish law.
- Four eyes: anything that writes (a decision, a rule, a report) goes to human approval.
- Self-hosted: training data generation and serving run on hardware we control.
3. Data
3.1 Source archive
The archive holds about 6,000 sources and 270,000 pages, split into passages that can be cited individually.
| Type | Share of sources |
|---|---|
| Guidance and supervisory publications | 61% |
| Academic articles | 17% |
| Theses | 7% |
| Laws | 5% |
| Regulations | 4% |
| Books | 3% |
| Communiqués and circulars | 2% |
The largest publishers are KVKK, MASAK, FATF, FFIEC, FinCEN, FINTRAC, the EBA, the Egmont Group, TBMM, the EDPB, the ICO, GİB and BDDK. Beyond Türkiye the archive covers the UK, the US, the EU, Germany, Canada, Montenegro, Croatia, Azerbaijan, Northern Cyprus, Switzerland, the UAE, Hong Kong, Australia, the Western Balkans and Central Asia. Turkish and English sources make up most of it; Montenegrin, Croatian, German and Azerbaijani sources are the fastest-growing part.
3.2 Domain coverage
| Domain | Typical sources |
|---|---|
| AML/compliance programmes | MASAK guidance, FATF Recommendations, Wolfsberg, AMLA technical standards |
| Fraud | BKM, TBB, TÖDEB, police cybercrime publications, Court of Cassation rulings |
| Data protection | KVKK Board decisions, EDPB guidelines |
| Transaction monitoring and analytics | Wolfsberg, MAS, HKMA and FCA monitoring guidance, calibration research |
| Tax and false documents | VDK and GİB guidance, rulings on fake invoices |
| Beneficial ownership, PEPs and corruption | trade registry rules, FATF and national guidance |
| Crypto assets, VASPs and the Travel Rule | SPK secondary regulation, MiCA technical standards, FATF updates |
| Suspicious transaction reporting and FIUs | MASAK reporting guides, FIU annual reports, Egmont |
| Terrorist financing and sanctions | OFAC, EU and UN list guidance, asset-freezing rules |
| Payments and e-money | TCMB, EBA, PSD3 |
| Illegal betting | court decisions, MASAK publications, academic work |
| Behaviour science and victim psychology | academic theses and articles, behavioural analytics research |
| Money mules, trafficking and organised crime | Europol, OSCE, UK Finance, Cifas, national warnings |
| Crime statistics | TÜİK, UNODC, Eurostat, Europol, national police and payment bodies |
3.3 Licensing
Every source is labelled either citable (laws, publications of public bodies, openly licensed material) or read-to-learn. Citable passages can appear in an answer with a reference. Read-to-learn material is shown only to the teacher models, never to TruvaLI verbatim, and every generated example is checked so it does not copy more than 15 consecutive words or 15% of its 8-word sequences from a source. Sources whose terms forbid AI training are excluded. Read-to-learn material is never translated; only the citable core is, because a translation is a derivative work.
3.4 Synthetic data generation
Training examples are written by open-source teacher models running on our own GPUs.
- Rejection sampling: the teacher writes four to eight candidates per task; the best goes into training, the worst valid one becomes the rejected half of a preference pair.
- Hard distractors: grounded questions are asked alongside similar but wrong passages, so the model learns to pick the right one.
- Unanswerable questions: a quarter of grounded questions have no answer in the given passages; the correct response is to say so.
- Behaviour simulator: 18 customer archetypes, about 30% of them legitimate, produce labelled behaviour scenarios so the model does not learn that every unusual customer is a criminal.
- Case cards: real cases reported in the news and in court decisions are turned into anonymised case cards, and each card is rewritten for the products and regulators of different sectors.
3.5 Quality gates
Every example passes an automatic validator before it can enter training:
- Schema and tool-call checks: structure, matching tool-call identifiers, valid JSON arguments
- Personal data patterns: Turkish ID numbers, IBANs, email addresses and phone numbers
- A tipping-off phrase list for any text addressed to a customer
- Rule vocabulary: the 168 fields and 53 aggregates of the TruvaLI AML rule language, operators and actions
- Case consistency: decision, flags and the reporting requirement must agree with each other and with the input figures
- Behaviour language: no clinical diagnosis, no suspicion from protected attributes, proportionate wording for victims
- Grounding: every citation must exist and every number must appear in the cited passage
- Copying limits and jurisdiction attribution
Near-duplicates are removed with MinHash at a 0.85 similarity threshold. Evaluation data is split off by customer, and 10% of the passages are held out entirely.
4. Task taxonomy
The design started from 30 core capabilities and has grown to more than 50 task codes.
| Family | Examples |
|---|---|
| Case work | case assessment, alert triage, STR narrative draft, four-eyes consistency |
| Rules and reports | rule proposal in the platform's rule language, calibration, report definitions, SQL |
| Knowledge | grounded regulation Q&A, typologies and red flags, exam questions, expert explanation |
| Screening | sanctions and PEP match assessment, Travel Rule, adverse media |
| Behaviour science | baselines, deviations, victimisation, mule life cycle, social engineering, transcripts |
| Platform and tools | MCP tool chains, structured output |
| Safety | refusals, prompt-injection resistance, insufficient evidence |
| Documents | ID and MRZ reading, business documents, forgery cues (planned) |
Answers are conditioned on nine roles: compliance, fraud, behaviour analysis, internal audit, psychology, tax inspection, risk, management, and fintech and crypto.
5. Training
5.1 Pilot (v0)
A small pilot model, trained on a few hundred examples, finished on 24 September 2026. It raised exam accuracy from 0.77 to 0.87 and grounded answers from 0.50 to 0.78, which was enough to justify the full run.
5.2 v1
v1 is a LoRA fine-tune of an open-source vision-language foundation model, trained in bf16 for one epoch. Loss was computed only on the model's own replies, and the vision encoder was left untouched, so TruvaLI keeps its base model's ability to read images. Thinking was not trained.
The v1 mix was dominated by knowledge: expert explanations (52%), exam questions (28%), grounded Q&A (8%), general behaviour and identity (5%), typologies (3%), social engineering (2%) and adverse media (2%). 80% of it was Turkish. That imbalance is deliberate for a first version and is the main thing v1.3 and v1.4 change: two thirds of the mix are skills such as tool use, case JSON, rules and SQL.
5.3 Serving
TruvaLI is quantised and served on our own server. PDFs are rendered page by page and transcribed by the model itself, so documents never leave the server. Thinking mode is the same model with reasoning switched on at inference time.
6. Evaluation
v1.1 against its untouched base model:
| Test | Questions | Base | TruvaLI v1.1 |
|---|---|---|---|
| Compliance exam (multiple choice) | 500 | 88.8% | 92.6% |
| Grounded Q&A, answer in the passage | 391 | 89.3% | 91.6% |
| Grounded Q&A, answer not in the passage | 9 | 77.8% | 100% |
| General behaviour and identity | 52 | 40.4% | 90.4% |
| Valid structured JSON output | 28 | 0% | 89.3% |
The exam pool has 17,976 questions; the grounded test draws on 803 held-out passages. The unanswerable and JSON results come from small samples and should be read as direction, not precision. From v1.4 on, every release is measured on larger sets: at least 50 unanswerable questions, at least 200 JSON tasks, 150 SQL questions in each of three dialects, and 50 questions per language for multilingual drift.
7. Safety
The core rules apply to every conversation: no tipping-off, no help with laundering or evasion, no decoding of masked data, embedded instructions treated as data, writes routed to four-eyes approval, and no invented figures. The behaviour tasks add limits of their own: no clinical diagnosis, no suspicion based on age, gender or nationality, and no claim to detect lies.
Dedicated safety and refusal training sets exist but were not part of v1; in v1.1 these rules are enforced through the system prompt. They are part of v1.3's training.
8. Limitations
- Tool use: v1 had no tool-use examples. Without thinking mode it failed all four tool-call tests; with thinking mode it passed all four.
- Figures: in testing it stated a reporting deadline of 24 to 48 hours where the regulation says ten business days. Figures and deadlines must be checked.
- Thinking was not trained: thinking mode works, but the reasoning itself was never part of the training data.
- Identity: without its system prompt the model does not always introduce itself correctly.
- Coverage: money mules, behaviour analysis, Turkish-language fraud and classic typologies are under-represented.
- Languages: the target-market languages make up 1% of v1's training data.
- Speed: responses are slower than we want on the current server.
9. Roadmap
The steps after v1.1 arrive as monthly minor releases until June 2027: nine of them, v1.2 to v1.10. None is a leap; each adds a part a compliance model should have. The first three ship before the end of this year.
Releases
| Release | Date | Main step |
|---|---|---|
| v1.2 | October 2026 | faster responses; a retrieval tool that looks up the source archive while answering |
| v1.3 | November 2026 | tool use and case assessment as JSON; safety and refusal training |
| v1.4 | December 2026 | rules and controls, SQL and reporting; preference training; 4,000 case cards; 3 sectors |
| v1.5 | January 2027 | 5 training languages; translated regulatory core for Montenegrin, Croatian, German and Azerbaijani |
| v1.6 | February 2027 | multilingual continued pre-training |
| v1.7 | March 2027 | anonymised case cards from our own closed cases; 6 sectors; 10,000 case cards |
| v1.8 | April 2027 | reinforcement learning with verifiers: SQL, rules and tool chains scored by code |
| v1.9 | May 2027 | document phase; 8 sectors |
| v1.10 | June 2027 | 10 languages; 10 sectors; 30,000 case cards; monthly data releases and quarterly retraining |
Before the end of this year: v1.2, v1.3 and v1.4
- Faster responses through speculative decoding
- A retrieval tool so TruvaLI can look up the source archive while it answers
- Skills become two thirds of the training mix: tool use, case assessment as JSON, rule and control design, SQL and report definitions
- The safety rules enforced today through the system prompt are taught to the model itself
- Preference training on pairs built from rejected candidates, from v1's own errors and from 👎 ratings
- 4,000 case cards from news, court decisions and published reports, each rewritten for banking, payments and e-money, and crypto-asset service providers
- An expert reading of 1,500 examples and a verified list of regulatory figures
By June 2027: v1.5 to v1.10
- Training in five languages: Turkish, English, Montenegrin and Croatian, German and Azerbaijani
- Continued pre-training on the native and translated archive, so TruvaLI reads and writes the market languages fluently and learns their own regulation
- Our own closed cases, anonymised under contract and a KVKK-compliant procedure, as case cards; raw records never enter training
- Six, then eight, then ten sectors: betting and gaming, insurance, currency exchange and precious metals, real estate, e-commerce and marketplaces
- Reinforcement learning where the answer can be checked by code: SQL results, rule test cases, tool-chain outcomes and regulatory figures
- A document phase: ID and MRZ reading, business documents, forgery cues and video-KYC support
- Ten languages and 30,000 case cards, with a quarter of training examples rooted in real cases
- After v1.10, monthly data releases, quarterly retraining and a yearly review of the foundation model
- A gold set reviewed and maintained by the TruvaLI AI community
On-premise
On-premise deployment does not wait for a release. Institutions that want to run TruvaLI on their own servers, alongside TruvaLI AML, can request it directly whenever they choose. See on-premise deployment for the details, or write to us to ask for it.
Release gates
Every release is measured on the same sets, which grow each time, and does not ship if it falls behind the release before it. These are the releases where the targets rise:
| Measure | v1.4 | v1.7 | v1.9 | v1.10 |
|---|---|---|---|---|
| Compliance exam | ≥92% | ≥93% | ≥93% | ≥94% |
| Grounded Q&A | ≥91% | ≥92% | ≥93% | ≥94% |
| Calls a tool where one is needed | ≥90% | ≥93% | ≥95% | ≥97% |
| SQL result accuracy, three dialects | ≥85% | ≥88% | ≥90% | ≥93% |
| Rules pass their test cases | ≥85% | ≥90% | ≥92% | ≥95% |
| Regulatory figure correct or flagged for checking | ≥95% | ≥97% | ≥97% | ≥98% |
| Multilingual drift, market languages | ≥85% | ≥90% | ≥92% | ≥94% |
| Agreement with the decision on our own closed cases | — | ≥75% | ≥80% | ≥85% |
10. Contributing
The roadmap depends on people who do compliance work for a living. If you can review answers, catch wrong figures, check a language or suggest sources, see the community page.