Software quality evidence is not the same thing as a testing programme, and the difference is what the first serious supervisory review tends to expose. Most regulated financial institutions have QA engineers, test suites, CI/CD pipelines and dashboards that show green. Far fewer have evidence.
A testing programme tells you whether software works. Evidence tells a supervisor, an auditor or a national competent authority that the right systems were tested, in the right way, at the right time, with the right oversight, and that the results were reviewed by people with the authority to confirm they were adequate. The cost of compliance in regulated industries sits mostly in that gap rather than in the testing itself.
Below we set out what the gap is, what it costs, why it persists, and what closing it actually requires.
What supervisors ask for that a test suite cannot answer
DORA moved from preparation into active supervisory review during 2026, which made this distinction visible in a way it was not before. National competent authorities are no longer checking whether institutions have a testing programme. They are checking whether the programme produces records that hold up when examined.
One request exposes the gap most clearly, and supervisors now make it in various forms: produce the resilience testing evidence for your three most critical ICT systems. Not the test results. A complete, traceable record of what was tested, why, when, by whom, with what outcome, and who confirmed the outcome was adequate. Chapter IV of Regulation (EU) 2022/2554, Articles 24 to 27, is where the obligation sits.
For institutions with mature evidence infrastructure, that request takes minutes. Where testing is well run but the evidence layer was assembled from CI/CD logs and spreadsheets, the same request triggers days of manual work. Results arrive incomplete and inconsistently formatted, which creates a second problem: a supervisor reading a hastily built package is now assessing evidence management practices alongside testing practices.
Why the software quality evidence gap exists
The gap is not negligence. It is the predictable result of how testing infrastructure evolved before regulatory evidence requirements existed.
Why manual evidence layers fail specifically
Manual processes that depend on individual discipline do not survive staff turnover, team restructuring or the compressed timelines of a supervisory request. They also fail quietly. Nobody notices that the sign-off spreadsheet stopped being updated in March until someone asks for the March records in October. That delay between failure and discovery is what makes the manual approach riskier than it looks on an org chart.
What the software quality evidence gap costs
The cost materialises in three ways, and none of them appears in a budget line labelled testing.
There is also a predictable date worth planning around. The Register of Information is submitted annually by 30 April, so at least one evidence exercise per year is on the calendar whether or not a review is scheduled.
The four requirements of software quality evidence infrastructure
Closing the gap is mostly an architectural decision about where evidence lives, who owns it and how it gets produced. Four requirements do the work.
A 90-day path to software quality evidence
Nobody rebuilds an evidence layer in a quarter. What fits in a quarter is enough structure to survive the next request, so the sequence below is ordered by what reduces exposure fastest rather than by what is most complete.
Weeks 8 to 12: retention and a dry run
Two habits make the difference afterwards. Rehearse the retrieval once a quarter, since an untested process is an assumption. And treat the annual Register of Information submission as the recurring deadline that keeps the chain current, because a layer maintained only when someone asks decays between requests.
Under pressure: what a major incident reveals
Nothing tests this infrastructure like a failure. A critical ICT system goes down, the response team activates, and the first decision is classification, because the 4-hour notification clock starts there.
Two institutions with the same test suite behave differently here. One has the record for the affected system, the most recent execution, the approving sign-off and the requirements covered in a single place within minutes, so the classification is better informed and the notification arrives complete. At the other, the same question sends people to CI/CD logs, email threads and spreadsheets while the clock runs. Their notification ends up reflecting the state of the evidence infrastructure rather than the state of the testing.
What software quality evidence does not solve
Worth being direct about the limits, because evidence infrastructure is sometimes sold as a compliance answer in itself.
A weak test suite stays weak. Perfectly traceable, immutably stored records of inadequate testing document the inadequacy with great precision, so the evidence layer improves the account of the work rather than the work. Deciding which systems support critical or important functions also remains an internal judgement, since that classification belongs to the institution's own risk assessment. Threat-led penetration testing keeps its own requirements and cadence for the entities in scope, and no evidence platform substitutes for it. Accountable humans stay in the loop too, because the sign-off is the part that cannot be automated by design. What does disappear is the assembly work, the inconsistency and the dependence on whoever happens to remember where things are.
How Qualigentic produces the software quality evidence chain
Qualigentic is an agentic QA platform built around the evidence layer rather than around test generation volume. Each test case carries the link to the requirement that motivated it. Execution records are timestamped and immutable. Sign-off runs through role-based workflows, so approval sits with people who hold the authority to give it. Export produces structured documents shaped for supervisory reading.
The distinction that matters: the evidence is produced continuously as testing runs, not assembled when someone asks. Deployment is on-premise, so records stay inside the institution's own perimeter, which also removes the third-party data transfer question. A time-boxed pilot covers one application and one framework and produces regulator-facing evidence in six to eight weeks. Our notes on AI QA in regulated industries and on the cost of software quality cover the surrounding arguments.


