One of the most common questions in quality management is:

How many transactions should we audit?

Should it be:

5 audits per employee?

10 audits?

1% of production?

100 transactions per month?

The answer is:

It depends on what you are trying to learn from the sample.

A sample designed to estimate the overall quality of a process may be very different from a sample designed to:

  • Coach an individual employee
  • Detect critical errors
  • Monitor a new process
  • Validate a policy change
  • Investigate an emerging defect
  • Review high-risk transactions
  • Compare teams
  • Meet a contractual or compliance requirement

This is why good QA sampling should not begin with a percentage.

It should begin with the purpose of the audit.

Why Sampling Exists

In many operations, auditing every transaction is impractical.

A processing team might complete:

10,000 transactions per month

or:

1,000,000 transactions per month

Auditing every transaction could require an enormous quality team and create little additional value.

Sampling allows the organization to review a manageable portion of the work and use that information to understand quality risk.

But a sample only works when it is designed carefully.

A poor sample can create false confidence.

An operation may report:

98% quality

while serious errors remain hidden because the audit sample rarely includes the transactions where those errors occur.

The First Question: What Are You Measuring?

Before selecting a sample size, determine what decision the quality result will support.

There are several common objectives.

Process-Level Quality

The organization wants to estimate the overall quality level of a process.

For example:

What percentage of all claims are being processed correctly?

This is primarily a statistical sampling problem.

Employee-Level Performance

The organization wants to understand whether an individual employee is applying the process correctly.

For example:

Does this processor require additional coaching?

The sample now needs enough coverage of that employee's work.

Risk Detection

The organization wants to identify critical or high-risk defects.

For example:

Are high-value transactions being processed incorrectly?

Pure random sampling may not be sufficient.

Improvement Analysis

The organization wants to understand which errors are driving quality loss.

The sample needs enough variation to support:

  • Error categorization
  • Pareto analysis
  • Root-cause analysis

Compliance Testing

A contractual, legal or regulatory requirement may define what needs to be tested.

In that situation, the organization's sampling methodology should follow the applicable requirement rather than a generic QA rule.

There Is No Universal “Correct” QA Percentage

Rules such as:

Audit 1% of production

sound simple.

But percentages can behave strangely as volumes change.

Suppose:

Team A processes 1,000 cases

Team B processes 100,000 cases

At 1% sampling:

Team A receives 10 audits

Team B receives 1,000 audits

Does Team B really require 100 times as many audits simply because it processes 100 times the volume?

Not necessarily.

Statistical sample requirements do not increase linearly with the size of the population.

This is why percentage-based sampling should be used carefully.

Statistical Sample Size: A Practical View

If the purpose is to estimate an overall quality percentage, sample size can be based on:

  • Confidence level
  • Margin of error
  • Expected defect rate
  • Population size

A commonly used conservative assumption is:

95% confidence

with:

±5% margin of error

and an assumed proportion of:

50%

The 50% assumption is used because it generally produces the largest required sample when the true quality level is unknown.

For a very large population, the required sample is roughly:

385 transactions

This surprises many operations teams.

Whether the population is 100,000 or 1,000,000, the sample does not need to become 1% of the entire population simply to estimate a proportion with that level of statistical precision.

For smaller populations, the required sample reduces because of finite-population correction.

Approximate examples are:

Monthly PopulationApprox. Sample at 95% Confidence / ±5%
500~218
1,000~278
5,000~357
10,000~370
100,000~383
Very large population~385

These numbers are useful for process-level estimation.

They should not automatically become employee-level QA targets.

That is a very important distinction.

Why Statistical Sampling Alone Is Not Enough for Operational QA

Imagine an operation processes:

100,000 transactions

and management audits approximately:

383 random transactions

This may provide a reasonable estimate of the overall defect rate under the assumptions above.

But suppose only:

0.2% of all transactions

belong to a particularly high-risk category.

A purely random sample might contain very few—or none—of those transactions.

The overall sample can therefore be statistically valid while still providing weak visibility into a specific operational risk.

This is why mature quality programs often combine:

Random Sampling

with:

Criteria-Based Sampling

and:

Risk-Based Sampling

Random Sampling

Random sampling gives transactions an equal or defined probability of being selected.

Its biggest advantage is reducing selection bias.

Without randomness, auditors or supervisors may unintentionally choose:

  • Easy transactions
  • Familiar cases
  • Recent work
  • Particular employees
  • Convenient transaction types

This can distort quality results.

Random sampling is especially useful for establishing a neutral baseline of overall process performance.

What Random Sampling Is Good For

Random sampling works well when management wants to understand:

What does normal production quality look like?

It is useful for:

  • Overall quality measurement
  • Trend analysis
  • Comparing periods
  • Establishing baseline performance
  • Reducing auditor selection bias

However, random sampling should not be expected to detect every rare or high-risk defect.

Criteria-Based Sampling

Criteria-based sampling intentionally targets transactions meeting particular conditions.

For example:

Audit transactions where:

  • Claim amount exceeds a threshold
  • Worktype equals Corrected Claims
  • Outcome equals Denied
  • Employee tenure is below 90 days
  • Product equals a particular category
  • AHT exceeds a defined limit
  • Transaction has been reworked
  • Customer complained
  • A particular error-prone process is involved

This is not truly random sampling.

And that is okay.

The objective is different.

The quality team is specifically investigating a defined population.

Risk-Based Sampling

Risk-based sampling takes this concept further.

Instead of treating every transaction as equally important, the QA program allocates more audit attention to areas where failures could create greater harm.

Risk may be based on factors such as:

  • Financial exposure
  • Customer impact
  • Regulatory impact
  • Compliance risk
  • Previous error rate
  • Transaction complexity
  • New process
  • New employees
  • Policy changes
  • Known control weaknesses

For example:

A ₹100 transaction and a ₹10,00,000 transaction may technically belong to the same process.

But an error on the second transaction may create considerably greater business exposure.

A risk-based audit program may therefore intentionally sample more heavily from the higher-risk population.

A Strong QA Program Often Uses Multiple Sample Types

Instead of choosing between random and targeted sampling, organizations can use both.

For example:

60% Random Sample

Provides a representative view of normal production.

25% Risk-Based Sample

Focuses on critical transactions.

15% Targeted Sample

Investigates known quality concerns.

This is only an illustration.

The correct split depends on the operation.

But the principle is important:

One sample can serve several different QA objectives.

Sampling New Employees

New employees often require a different sampling approach.

A new processor may need:

  • More frequent audits
  • Larger samples
  • More immediate feedback
  • Broader Worktype coverage

As confidence in their performance increases, audit frequency may decrease.

For example, an organization might define stages such as:

Training

High audit coverage

↓

Early Production

Enhanced sampling

↓

Established Performance

Standard sampling

↓

Sustained Strong Performance

Reduced routine sampling with continued risk-based coverage

The exact rules should reflect the organization's risk tolerance.

Should Low Performers Receive More Audits?

Potentially, yes.

Suppose an employee repeatedly falls below the quality target.

Increasing audit frequency temporarily can help answer:

Is the issue isolated?

or:

Is there a persistent performance problem?

Additional sampling can also provide enough evidence to identify recurring error categories.

However, the purpose should be improvement—not simply generating more failures.

The workflow should connect:

Higher sampling

↓

Better diagnosis

↓

Targeted coaching

↓

Follow-up audit

↓

Improvement validation

Strong Performers Should Not Disappear from QA

A common mistake is to stop auditing high performers almost completely.

That creates blind spots.

Performance can change because of:

  • New policies
  • New Worktypes
  • Increased workload
  • Role changes
  • Process complexity
  • System changes

High performers may receive lower audit intensity, but some continuing random coverage is usually valuable.

Sample Across Different Worktypes

Suppose an employee processes:

60% Worktype A

30% Worktype B

10% Worktype C

If all audits come from Worktype A, the employee's overall score may not reflect their complete workload.

A quality program may therefore consider stratified sampling.

This means dividing the population into meaningful groups and sampling from each.

For example:

Professional Claims

Hospital Claims

Corrected Claims

Then selecting transactions from each group.

This can provide stronger coverage when quality risk varies across Worktypes.

Stratified Sampling Can Improve Representation

Imagine monthly production consists of:

Professional Claims — 70%

Hospital Claims — 20%

Corrected Claims — 10%

A purely random sample will probably follow a similar distribution.

That may be appropriate for estimating overall quality.

But suppose Corrected Claims have significantly higher risk.

The QA team may intentionally oversample that Worktype.

For example:

Professional — 55%

Hospital — 25%

Corrected — 20%

The resulting sample no longer represents the production mix exactly.

But it provides better visibility into the risk area.

The reporting should clearly distinguish between:

Representative quality measurement

and:

Targeted quality monitoring

Avoid Mixing Targeted Audits into the Headline Score Without Thought

This is important.

Suppose the QA team intentionally selects transactions that are more likely to contain defects.

If those targeted audits are combined directly with the random sample, the resulting quality score may appear worse than the true overall population.

Conversely, selecting only easy transactions may make quality appear artificially high.

Organizations should therefore clearly define how different sampling streams contribute to reported metrics.

For example:

Random QA Score

for representative process quality

and:

Targeted Risk Findings

for investigative monitoring

may sometimes be more meaningful than blending everything into one number.

Sampling Should Cover the Entire Period

Another common issue is timing.

Suppose a monthly QA target requires:

10 audits per employee

All ten audits are completed during the final three days of the month.

Technically, the target was achieved.

Operationally, the approach is weak.

If the employee had been making the same error throughout the month, the organization lost several weeks of opportunity to correct it.

Sampling should ideally be distributed throughout the operating period.

That enables:

Audit

↓

Feedback

↓

Correction

↓

Measure improvement

rather than discovering all problems at month end.

Feedback Speed Matters as Much as Sample Size

Auditing 50 transactions per employee may provide little value if feedback arrives six weeks later.

A smaller sample with timely feedback may create more improvement than a much larger delayed sample.

Quality programs should therefore monitor not just:

How many audits were completed?

but also:

How quickly were findings communicated?

This is where the QA workflow becomes as important as the sampling methodology.

Sampling Should Consider Critical Errors

Some failures are rare but severe.

For example:

  • Regulatory violation
  • Incorrect financial decision
  • Privacy breach
  • Wrong customer outcome
  • Critical safety issue
  • Major compliance failure

If a critical error occurs only once in every 1,000 transactions, a small random sample may easily miss it.

Organizations may therefore need additional controls such as:

  • Targeted sampling
  • Automated validation
  • Full-population rules
  • Exception reporting
  • High-risk transaction review

Sampling is not always the correct control for every risk.

Audit Capacity Still Matters

Statistical theory is useful.

But quality teams operate with finite resources.

Suppose the recommended sampling design would require:

5,000 audits per month

but the QA team has capacity for:

2,500

The organization needs to make choices.

It may prioritize:

  1. Critical-risk coverage
  2. Minimum employee coverage
  3. Random process sampling
  4. Targeted investigations

The best QA model is therefore a balance between:

Statistical confidence

Business risk

and:

Available audit capacity

Calculate QA Capacity Before Setting Targets

Suppose one audit takes:

15 minutes

An auditor has:

6 productive audit hours per day

Each auditor can theoretically complete:

6 × 60 ÷ 15 = 24 audits per day

With:

20 working days

monthly theoretical capacity is:

24 × 20 = 480 audits

per auditor.

With:

5 auditors

theoretical monthly capacity becomes:

2,400 audits

Before setting sampling requirements, management should understand whether the QA team can realistically perform them.

Otherwise, sampling targets can create:

  • End-of-month rush
  • Superficial audits
  • Delayed feedback
  • Auditor burnout
  • Reduced audit quality

Audit Quality Is More Important Than Audit Quantity

Increasing the sample does not automatically improve the QA program.

Consider:

1,000 poorly executed audits

versus:

500 well-selected, consistent audits with timely feedback and root-cause analysis

The second program may create substantially more improvement.

Quality management should therefore balance:

Coverage

with:

Audit consistency

Feedback quality

Calibration

Actionability

Calibration Is Essential

Sampling determines what gets audited.

Calibration helps ensure auditors judge those transactions consistently.

If two auditors review the same case and reach different conclusions, increasing the sample will not solve the underlying problem.

Organizations need clear definitions for:

  • Audit questions
  • Scoring
  • Error categories
  • Severity
  • Critical errors
  • Not-applicable responses

Regular calibration improves the reliability of the QA result.

Sample Size Should Consider Employee Count

There is another operational issue.

Suppose a process requires approximately 385 audits for statistical process-level measurement.

But the team has:

200 employees

If those audits are distributed randomly, some employees may receive only one or two audits—or none.

That may be acceptable for estimating overall process quality.

It is not enough for reliable individual performance management.

This is why QA programs often need two different layers:

Process-Level Sample

Designed to estimate overall quality.

Employee-Level Minimum Coverage

Designed to provide enough observations for coaching and performance management.

The two requirements should not be confused.

A Practical Hybrid Sampling Framework

For many operational QA programs, a practical approach might include four layers.

Layer 1 — Baseline Random Sampling

Measure normal process quality without intentional bias.

Layer 2 — Minimum Employee Coverage

Ensure every eligible employee receives some regular audit coverage.

Layer 3 — Risk-Based Sampling

Increase coverage of high-risk transactions, Worktypes or outcomes.

Layer 4 — Targeted Sampling

Investigate known issues, new processes, repeat errors or specific concerns.

This creates a much stronger quality framework than using one fixed percentage for everything.

Example of a Monthly Sampling Design

Suppose an operation processes:

50,000 transactions per month

with:

40 processors

and:

5 QA auditors.

Management might design separate sampling streams.

For example:

Baseline Random Sample

Used to understand overall process quality.

Employee Minimum Sample

Used to ensure each processor receives appropriate QA coverage.

High-Risk Sample

Used for specific critical transaction categories.

Targeted Improvement Sample

Used for employees or error types requiring deeper review.

The exact numbers should be based on:

  • Desired statistical precision
  • Business risk
  • Available audit capacity
  • Employee population
  • Historical performance

This is more defensible than simply saying:

“We audit 2% of everything.”

Sampling Strategy Should Change Over Time

QA sampling should not be configured once and forgotten.

Suppose a particular error type increases significantly.

Sampling may need to shift toward that area.

Suppose a Worktype consistently achieves very strong quality with low risk.

Routine sampling could potentially be reduced.

Suppose a new product launches.

Audit intensity may need to increase temporarily.

A mature sampling strategy responds to what the quality data is showing.

Use Pareto Analysis to Influence Future Sampling

Suppose Pareto analysis reveals:

Error A — 35% of defects

Error B — 27%

Error C — 18%

Together, those three categories create:

80% of observed defects

That information can influence future targeted sampling.

Quality management becomes cyclical:

Sample

↓

Audit

↓

Identify Error Pattern

↓

Adjust Sampling

↓

Coach / Correct Process

↓

Audit Again

↓

Measure Improvement

Sampling therefore becomes part of the improvement strategy rather than an administrative target.

Avoid Over-Sampling the Same Employees

Criteria-based sampling can accidentally create unfair audit concentration.

For example, if one employee handles most high-value transactions, risk-based sampling may repeatedly select that person's work.

Their audit count may become significantly higher than their peers.

That is not automatically wrong.

But management should understand why it is happening.

Reporting should distinguish between:

Audit volume

and:

quality performance

A higher number of audits does not automatically mean someone is performing poorly.

Track the Sample Against the Population

A mature QA program should understand both sides:

Production Population

What work was actually completed?

Audit Sample

What part of that work was reviewed?

This allows management to compare:

  • Worktype mix
  • Employee mix
  • Product mix
  • Outcome mix
  • Complexity
  • Risk categories

If the sample looks dramatically different from production without a deliberate reason, the QA result may require careful interpretation.

What Should a QA Sampling Dashboard Show?

Useful sampling information may include:

Total Production

Total Audits

Audit Rate

Employees Covered

Worktypes Covered

Random Sample Count

Criteria-Based Sample Count

High-Risk Sample Count

Audits Pending

Audits Completed

Quality Score

Critical Error Rate

Top Error Categories

Repeat Errors

This provides visibility into both the quality result and the methodology behind it.

Questions to Ask About Your Current Sampling Program

A quality manager should be able to answer:

Why did we choose this sample size?

What decisions is the sample intended to support?

Is the sample representative of production?

Are high-risk transactions receiving enough coverage?

Does every employee receive appropriate audit coverage?

Are new employees sampled differently?

Are low performers sampled differently?

Are audits distributed throughout the month?

Are targeted audits being mixed into the headline score?

Do we have enough auditor capacity?

Does sampling change when risk changes?

If the answer to most of these is:

“We have always audited this number.”

the sampling strategy may be ready for review.

A Better Way to Think About QA Sampling

Instead of asking:

“What percentage should we audit?”

ask:

“What level of evidence do we need to make the decision?”

Then consider:

Statistical confidence

Employee coverage

Risk

Process complexity

Audit capacity

That creates a much more purposeful sampling framework.

How Praevexa QualityFlow Supports QA Sampling

Praevexa QualityFlow is designed to help quality teams create structured sampling and audit workflows rather than manually selecting transactions from spreadsheets.

QualityFlow supports capabilities including:

  • Production data upload
  • Configurable QA Worktypes
  • Minimum sample rules
  • Random sampling
  • Criteria-based sampling
  • Multiple categorical and numeric sampling criteria
  • Auditor skill and ownership controls
  • Configurable audit checklists
  • Error categories and severity
  • Audit scoring
  • Processor feedback and acknowledgement
  • Disputes and manager resolution
  • Pareto analysis
  • Quality trends
  • Hierarchy-based performance analysis
  • Excel reporting and data exports

The objective is not simply to increase the number of audits.

It is to help organizations create better audit coverage, clearer quality evidence and more actionable improvement data.

Learn more about Praevexa QualityFlow:

https://www.praevexa.com/QualityFlow.aspx

Related Reading

What Should a Quality Management System Actually Do? 12 Capabilities Beyond QA Scoring

https://www.praevexa.com/insights/quality-management-system-essential-capabilities

Why Quality Scores Alone Don’t Tell You Where the Process Is Failing

https://www.praevexa.com/insights/why-quality-scores-alone-are-not-enough