Context and problem
Before signing a contract or paying an advance, a business wants to know whether the other company is reliable. In Russia most of the signals are public: the state register of companies, filed financial statements, arbitration court cases, enforcement proceedings and public procurement contracts.
They live in different places and formats, and reading them takes time and expertise. The question itself is vague: “is it safe to work with them?” is not a field in any database.
Understanding
People did not need more raw data. They needed a decision aid: a clear first signal, the reasons behind it and a short list of what to check before the deal.
They also needed to trust it. A confident AI text that cannot be traced back to data is worse than no answer, so “no data” must never be presented as “no problems”.
MVP and scope
- In
- One input — a company tax ID (INN). A limited set of public sources through one data aggregator. One report: verdict, reasons, confirmed facts, risks and open questions, limitations, and what to check before the deal. Abuse protection and daily limits, so the public demo cannot be used as a free API.
- Out
- User accounts, billing and other commercial features. Full company dossiers with every available source. Monitoring companies over time.
What I built
Open the demo and a saved report on a real company is already on the screen — no tax ID needed. Five saved examples cover green, yellow and red. Enter a tax ID and the service fetches fresh data, shows the verdict first and then streams the written report.
Behind the screen, the API collects data from the aggregator within strict time and size limits and compresses it to what matters for the assessment. A rule-based engine computes the verdict, one LLM call writes the explanation, and a deterministic check rejects text that contradicts the data on key points.
What you will see: the interface and the reports are in Russian. The traffic light, the list of sources and the structure of the report are readable without it.
My ownership
- Built by me
- The original RISK product: the product concept and main workflow, the core business logic, data collection through the Checko aggregator and work with raw company data, LLM request construction, the AI analysis and generation flow, the main user scenario and the original UI/UX of the public RISK.
- Reused later
- The commercial product Legicon later reused this core.
- Another developer
- In the commercial phase another developer reworked the UI and layout and added registration and accounts, email, tariffs and billing, and payments.
- Current demo
- The public demo runs on the same core, carried over as a separate package. The demo layer around it — the traffic light, saved examples and abuse protection — I built with AI coding agents: I made the product decisions, reviewed the code, and ran the tests and releases.
Key decisions
- The engine sets the colour; the LLM explains it
- The alternative was to let the model judge. A verdict has to be reproducible and auditable, and a language model is good at explaining, not at being consistent. The model receives the computed verdict and is not allowed to change it.
- A real saved report on the first screen
- Most visitors have no tax ID at hand, and an empty form shows nothing. A saved report costs nothing to open and shows the result immediately. Of five prototyped first screens, I shipped a different one than the agent recommended.
- Self-hosted proof of work instead of a third-party captcha
- The external captcha was checked only in the browser, so the protection failed open. The replacement is ALTCHA proof of work inside the existing API: one-time proofs, atomic counters in SQLite and no new services.
- Carry the original core over instead of rewriting it
- A separate TypeScript port existed but had no proof of parity. The original Python core was moved into a package unchanged, checked against a SHA-256 manifest, with its own tests.
- Keep the stronger model, cut cost elsewhere
- A cheaper model returned an unusable reasoning-only answer in the comparison. The savings came from a shorter demo prompt and fewer data sources instead.
Constraints, mistakes and trade-offs
- Compression is lossy: the model sees a compact summary, not every record. Unknown, missing and zero values stay distinct so that missing data does not read as a clean record.
- The demo uses five data calls for a company; the full profile behind the saved examples needs up to sixteen. More sources mean more cost and latency per check.
- A fresh check takes about 15–30 seconds, most of it the LLM. The verdict appears first so the wait is not empty.
- A live smoke test showed the model echoing English field names into the Russian report. Unit tests could not catch it; switching the model’s input to Russian keys fixed it.
- Saved examples are dated snapshots from 1–2 October 2026, not live data.
Validation and testing
- The core carries about 200 unit tests, plus frozen capacity cases for prompt size.
- Releases pass API, frontend and release-guard suites; the 1 October 2026 release ran 248 unit, 107 integration, 22 release-guard and 27 UI tests.
- Saved examples pass an offline gate with no network: every sum, date and case number is matched to the source data, and the factual check and colour guard run before publishing.
- The examples are not raw AI output: the green ones are model drafts corrected by hand, and the yellow and red ones were written from the sources.
- Each release is checked in a browser at desktop and mobile widths on the exact build, then with a live check in production.
Deployment and real status
RISK has been public since March 2026; the current version went live in October 2026. Builds happen on a separate machine, and production receives the exact tested container image.
Limits: 20 new checks per IP address and 100 in total per day. Russian companies only, interface in Russian. This is a public demo, not a commercial service.
What I would improve next
- Measure the proof-of-work solve time on real phones and tune it.
- Link each statement in the report to the exact data it came from, inline.
- Refresh the saved examples on a schedule, so their dates do not age.
Tech summary
- Frontend
- Nuxt 3, Vue 3, TypeScript; a static build served by nginx
- API
- Python, FastAPI, streaming over server-sent events, SQLite for limits and one-time proofs
- Core
- Dependency-free Python package: rules, data compression, prompts, output validation
- AI
- One streamed LLM call per check via OpenRouter; deterministic verdict and output checks
- Data
- Checko aggregator over Russian public registers
- Protection
- ALTCHA proof of work, per-IP and global daily limits
- Infrastructure
- Docker Compose behind nginx; build and production on separate servers