All writing
Audit Methodology6 August 20266 min read

Sampling Was a Constraint, Not a Standard

The theory of sampling is a rigorous answer to a physical limitation. The limitation is the reason — not the other way round.

In short
  • Sampling under SA 530 is a rigorous response to a physical limit: a person cannot read eleven thousand vouchers.
  • Sampling misses patterns spread across items, such as structured cash payments or duplicate payments.
  • For tests on the recorded data, full-population testing replaces sampling; for evidence outside the books, sampling remains.

Ask a group of chartered accountants why we sample, and you'll get a reasonable answer about SA 530, about representative selection, about materiality and assurance thresholds and the impossibility of absolute certainty.

All true. But it's the second answer. The first answer, the one underneath it, is simpler: because a person cannot read eleven thousand vouchers.

Sampling is a beautiful, rigorous, well-theorised response to a physical limitation. The theory exists because the limitation existed. And it's worth sitting with the fact that the limitation is the reason — not the other way round.

What the sample structurally cannot see

Sampling is very good at estimating how often something happens. If 3% of your vouchers have a documentation defect, a well-drawn sample will find that out.

It is systematically bad at one particular thing: anything whose meaning lives in the pattern rather than the item.

Take a payment of ₹9,800 in cash. Pull it in a sample and it's clean. Below the section 40A(3) limit of ₹10,000, properly vouched, nothing to report. Move on.

Now look at the whole population. That same vendor received ₹9,800 in cash on the 3rd, the 11th, the 19th and the 27th of the same month. Four payments, one day apart from being obviously structured, every one of them individually unremarkable. The sample saw one of them and found nothing wrong, and the sample was right about the item it saw. It just wasn't the item that mattered.

The same shape recurs everywhere once you start looking:

  • A vendor invoiced at ₹49,500 eleven months running, sitting just under a TDS threshold. Any one invoice is fine. Eleven of them is a fact about the relationship.
  • Two payments to the same party, same amount, same week, different voucher numbers. Neither is unusual alone. Together they're a duplicate payment, and the money has left.
  • A cash balance that dips negative for four days in November and recovers. No individual voucher is defective. The sequence is impossible.
  • An expense head with forty entries, thirty-eight of which are round thousands. Nothing wrong with a round number. Thirty-eight of them is a pattern in how somebody is entering data.

None of these are findings you get to by examining a well-chosen item more carefully. They're findings that only exist at the level of the population, which is precisely the level sampling was designed to let you avoid.

The other thing sampling was hiding

There's a second cost, less discussed. Sampling doesn't just limit what you find — it limits what you can say.

When a client pushes back on a disallowance, the strength of your position depends on what you can put in front of them. "We tested a sample and extrapolated" is a defensible professional position and a weak conversational one. "Here are all 214 transactions with this party across the year, here are the 31 that fell foul of the provision, and here is the voucher for each" ends the conversation.

Same for peer review. Same for a departmental query eighteen months later, when the person who did the fieldwork has left the firm and the file has to speak for itself.

What actually changes

If a machine reads every voucher — not a sample, the population — the whole shape of the engagement moves.

This is how Audcrix is built. It reads vouchers and ledgers directly from Tally, reconciles what it read against Tally's own trial balance and P&L, and stops rather than proceeding if those don't agree. Then it tests all of it.

Ledgers get classified by structure, not by name. Grouping by keyword breaks the moment a client names an account something unexpected — and clients always do. Classification by the actual group chain in Tally means "Sundry Expenses A/c No. 2" lands where it belongs regardless of what somebody typed.

Thresholds are applied to every transaction. Not to the ones that made the sample. Section 40A(3) cash payments, section 43B(h) MSME payments beyond the prescribed period, section 14A disallowances — every voucher is measured against the applicable limit, and every add-back arrives with its section and its supporting vouchers attached.

TDS is matched to the actual payee. This one is quietly important. Firms routinely test TDS by looking at expense heads, which works right up until the payee behind the expense isn't who the head implies. Resolving the real party from the voucher and then matching deduction by deductee and by section catches the mismatches that head-level testing is structurally blind to.

26AS gets reconciled party by party. Not totalled and compared — matched, with the specific gaps named. Which party, which amount, which direction.

Twelve months of GST are matched against what was actually filed. Books against GSTR-1, 2A, 2B and 3B, period by period, rather than a year-end totals check that nets off two errors in opposite directions and reports nothing.

Forensic tests run across everything, because they only work across everything: negative cash balances, duplicate payments, Benford anomalies, suspicious sequences.

Related parties get mapped once and applied consistently. The AS 18 failure mode has never been that firms don't know the related parties. It's that the mapping lives in one person's head, and gets applied to the transactions that person happened to look at. Map once, apply everywhere, and the disclosure is complete by construction rather than by diligence.

Does the standard still matter?

Yes — and this is worth being precise about, because "we test everything now" can be heard as "we've stopped doing audit methodology," which would be the wrong lesson.

SA 530 isn't repealed by better tooling. What changes is where sampling is necessary versus where it's a choice. For anything that can be computed from the recorded data — arithmetic, thresholds, matching, classification, sequence — there is no longer a reason to sample, and the working papers on materiality, risk assessment, sampling approach and analytics should say what you actually did.

For everything that requires going outside the books, sampling remains exactly what it was. Physical verification. Confirmations. Inspecting a contract. Assessing whether a transaction has commercial substance. Those still need selection, and they still need judgement about what to select.

The distinction is clean: testing the population replaces sampling for data; it doesn't replace sampling for evidence that lives outside the data. A firm that understands that line is doing better methodology after the change than before it, not worse.

The honest version of the argument

I don't want to oversell this. Full-population testing doesn't make an audit correct. It doesn't tell you whether a provision is adequate or whether management is telling you the truth. It won't form your opinion.

What it does is remove a specific, known, decades-old blind spot — the one where a defect is invisible because it's distributed across items that are each individually fine — and hand you the exceptions instead of the ledger.

You still have to look at them. You still have to decide what they mean. That was always the job.

It's just that for a long time, the job came bundled with a physical constraint that shaped how we thought about assurance. The constraint is gone. It seems worth updating the habit that grew around it.

Questions this answers

Does full-population testing replace SA 530 sampling?

For anything computed from the recorded data, such as arithmetic, thresholds, matching, classification and sequence, there is no longer a reason to sample. For evidence outside the books, such as physical verification, confirmations and inspecting a contract, sampling remains.

What can audit sampling not detect?

Anything whose meaning lives in the pattern rather than the item: four cash payments of ₹9,800 to one vendor in a month, invoices parked just under a threshold, duplicate payments under different voucher numbers, or a cash balance that goes negative for four days.

Why do auditors sample at all?

Because a person cannot read eleven thousand vouchers. Sampling is a rigorous, well-theorised response to that physical limitation.

Does testing every voucher make an audit correct?

No. It does not tell you whether a provision is adequate or whether management is telling the truth. It removes the blind spot where a defect is spread across items that are each individually fine.

Where Audcrix runs this

  • Fraud DetectionForensic checks across every Tally voucher — duplicate payments, negative cash, Benford anomalies, back-dated postings — with convergence and honest gap flags.
  • Statutory AuditStatutory audit software built around the Standards on Auditing — materiality, SA 530 sampling, SA 510 opening balances and confirmations, documented as the work is done.