How we calculate the tunneling risk score
Every company on Nanayojana carries a tunneling risk score — a single number that summarises patterns across years of annual reports associated, in academic and regulatory literature, with controlling shareholders or insiders extracting value from a company at the expense of minority shareholders.
This page explains what goes into that score and how it's calculated. Some implementation details are intentionally withheld — the reasoning is explained at the bottom of this page.
Data flow
The flow, end to end
Every score is built from public annual report filings — nothing else. There is no manual input, no analyst opinion, and no data purchased from third parties.
Read the report
Each annual report is parsed section by section — chairman statement, MD&A, notes to accounts, related-party disclosures, auditor's report, and more.
Understand the language
An NLP sentiment model reads each section independently, so an upbeat chairman's letter can be weighed against a more guarded set of notes in the same report.
Extract the relationships
Directors, subsidiaries, counterparties, and the transactions between them are extracted and linked into a company-and-people network.
Detect & score
Defined patterns across sentiment, transactions, and governance are checked against every report. Each match becomes a recorded, evidenced signal that feeds the score.
Signal detection
What we look for
Every detected signal falls into one of three broad evidence themes. Each one is recorded against a specific company, year, and report section — you can see the underlying evidence for any company on its company page. The complete, filterable list of individual signals — with points, live counts, and examples — is in the signal catalogue below.
Sentiment & narrative
We compare tone across sections of the same report. A chairman's statement or outlook section that reads confidently positive, sitting alongside notes, audit commentary, or related-party disclosures that read materially more negative, is a divergence worth flagging — the public narrative and the fine print are telling different stories. Persistently negative language in the risk-bearing sections of a report is scored on its own, even without divergence.
Related-party transactions
Related-party transactions are the primary channel through which value is tunneled out of a company, so we pay close attention to them. Signals in this category capture things like: transactions disclosed without a stated value, transactions that are unusually large relative to the company, a related-party footprint that looks disclosed on paper but doesn't square with the entity relationships we extract elsewhere, transaction activity unusually light compared to similar companies in the same sector, and related-party relationships that quietly stop being disclosed after appearing in prior years.
Governance & audit
These capture weaknesses in the checks that are supposed to catch tunneling before it happens: a qualified audit opinion or emphasis of matter, and directors who appear on both sides of a transaction — sitting on the board while also being a named counterparty to a related-party deal.
We publish each signal's raw point value in the catalogue below, but deliberately don't publish the exact detection thresholds or currency cutoffs that make it fire — see "why we don't publish the exact formula" below.
Signal catalogue
Every signal the model can emit
The complete list of risk signals, grouped by the five weighted categories that make up the score. Points are the raw value each signal carries per occurrence, before category weighting and time decay. Counts and examples are live from the evidence database.
Sentiment
Weight: 15% of the final scoreTone divergence between public and risk sections, negative risk sentiment, obfuscation.
| Signal | Points | What it means | Detections | Latest example |
|---|---|---|---|---|
|
Sentiment divergence
sentiment_divergence
|
+2 pts | Public-facing sections (chairman's statement, outlook) read confidently positive while risk-bearing sections of the same report read negative — narrative and fine print tell different stories. | 0 detections | — |
|
Negative tone in risk sections
negative_risk_sentiment
|
+1 pts | A risk-bearing section (notes, auditor's report, related-party disclosures) reads strongly negative on its own. | 1 detection | MERC.N0000 · FY2015 |
|
Negative sentiment jumped year over year
yoy_negative_sentiment_spike
|
+2 pts | Negative sentiment in a risk-bearing section rose sharply compared with the previous year's report. | 2 detections | MERC.N0000 · FY2015 |
|
Hard-to-read risk section
section_obfuscation
|
+1 pts | A risk-bearing section is written in unusually complex language and reads negative — a combination associated with obfuscation. | 2 detections | MERC.N0000 · FY2015 |
|
Dense or evasive disclosure language
high_complexity_disclosure
|
+1 pts | A risk-bearing section scores high on a composite of long sentences, long words, passive-voice phrasing, and accounting jargon — independent of tone, this style pattern is associated with obscuring rather than clarifying disclosure. | 30 detections | DIAL.N0000 · FY2019 |
|
Elevated uncertainty or litigious language
elevated_uncertainty_language
|
+1 pts | A risk-bearing section uses uncertainty words (e.g. "contingent", "fluctuate", "unforeseen") well above typical frequency alongside a negative tone, or uses litigation-related language (e.g. "claimant", "lawsuit", "arbitration") at an unusually high rate. | 89 detections | PKME.N0000 · FY2024 |
|
Unchanged disclosure year over year
boilerplate_disclosure
|
+1 pts | A risk-bearing section reads almost word-for-word the same as last year's report — text that barely changes can mean the underlying situation genuinely hasn't changed, or that disclosure isn't keeping pace with it. | 31 detections | SWAD.N0000 · FY2018 |
No signals match your search.
Score mechanics
Two things that shape the final number
Recency matters
Risk only ever accumulates from detected signals — nothing is ever manually subtracted. But older signals count for progressively less than recent ones, because ownership and control structures at Sri Lankan listed companies tend to change slowly, and a pattern from many years ago is weaker evidence of current risk than one from last year's report. A company with a clean, multi-year track record since an old signal will see its score fall accordingly, even though the historical record is never erased.
Networks matter
Tunneling is rarely contained within one company — it typically runs through a director, a parent, a subsidiary, or a recurring counterparty. So a company's score isn't computed in isolation: some risk from the companies it's connected to, through shared directors or transaction relationships, spills over. A company with no signals of its own can still carry an elevated score if its directors or counterparties are strongly associated with higher-risk companies. This spillover is always smaller than owning the risk directly.
Result
The final score and risk tiers
A company's own accumulated, time-decayed signals and the risk absorbed from its network are added together into one tunneling risk score. To make the number easier to scan, every company is also placed into a risk tier shown as a colour-coded badge throughout the site.
The numeric boundaries between tiers aren't published, for the same reason the detection thresholds and decay mechanics aren't: a company motivated to manage its score rather than its underlying risk could otherwise structure disclosures to land just under a line.
Every point traces back to evidence
Nothing in the score is a black box at the level of an individual company. Open any company's network page and you can see every signal that contributed to its score — the year it was detected, the report section it came from, and the extracted evidence text, with a link back to the original passage in the source annual report. What isn't published is the general recipe that would let the pattern be replicated or gamed across every company at once.
Read this before you trust a number
What this score is — and isn't
A pattern-detection tool built from public annual report disclosures, not an accusation of wrongdoing. Every signal has an innocent explanation in some cases — a large related-party transaction may be entirely legitimate and properly governed.
Not investment advice and should not be the sole basis for any investment decision.
Dependent on disclosure quality. A company that discloses more is, everything else equal, more likely to trigger a signal than one that discloses less — we try to offset this with disclosure-gap signals, but the score can only see what's in the filed report.
Sentiment analysis is imperfect. Language models can misread hedged, technical, or translated text, which is why sentiment signals are only one of three categories and are weighed against corroborating evidence elsewhere.
Recalculated periodically, as new reports are processed, not continuously in real time.
This isn't a novel theory — it's applied academic research
Tunneling, the Entrenchment Effect, information asymmetry in low-float markets, and textual tone as a predictor of returns are all established findings in finance and governance literature. See the cited sources behind this approach, including Sri Lanka-specific corporate governance research.
Why we don't publish the exact formula
We've deliberately published the categories of evidence we look at, the individual signals and their raw point values, and the mechanics that shape the score — recency, network spillover, and additive accumulation — while withholding the exact detection thresholds, currency cutoffs, decay rate, spillover weight, and tier boundaries. This is a standard trade-off for any risk or credit-style score published alongside the entities it scores (rating agencies and credit bureaus draw the same line): full transparency down to the formula would let a company reverse-engineer the minimum disclosure needed to stay under a threshold, which would optimise for a better score rather than genuinely lower risk — and it would let the same recipe be copied wholesale by anyone who hasn't done the underlying data work. Every individual finding remains fully auditable on a per-company basis; it's the general recipe, not the evidence, that stays behind the curtain.