Treatment integrity measures whether an intervention was delivered as designed. Interobserver agreement, or IOA, measures whether independent observers recorded the same events or values. Clinicians usually need both when treatment decisions depend on directly observed implementation and client data. High scores answer different questions: integrity supports confidence in what was delivered, while IOA supports confidence in how consistently it was measured. Effectiveness, fit, ethics, and social validity still require separate evidence.
The short comparison: treatment integrity vs IOA
QuestionTreatment integrityIOAWhat does it evaluate?Correspondence between the written or trained procedure and actual implementationCorrespondence between two independent records of the same observationWhat is being measured?Implementer behavior, treatment components, timing, dosage, or exposureObserver records of client behavior, implementer behavior, or another defined eventWhat can a strong result support?Confidence that the intended procedure occurred at the measured levelConfidence that trained observers applied the measurement system similarly in the sampled observationsWhat can it miss?Observer error, a weak plan, poor contextual fit, unwanted effects, or a critical component hidden by an averageShared observer bias, an invalid definition, a poor sampling plan, or incorrect implementationTypical response to a weak resultExamine component errors, feasibility, training, resources, plan fit, and environmental barriersRepair definitions, observation training, independence, sampling, matching rules, and calculation choice
A treatment integrity vs IOA review works best when the clinician adds a third signal: the client's outcome data and experience. The three measures answer whether the plan was delivered, whether the observations were dependable enough for the decision, and whether the plan helped in a way that mattered to the client.
Define the measures before the observation
For treatment integrity, convert the active plan into observable implementation components. Each component needs:
- a clear start and end
- the conditions that create an opportunity
- the response that counts as correct
- omission and commission errors when each matters
- permitted adaptations, prompts, and response forms
- a rule for unavailable or invalid opportunities
- critical safety, assent, communication, and dignity safeguards
- the person responsible for implementation
A common global calculation is:
correctly implemented component opportunities ÷ applicable component opportunities × 100
Also calculate each component separately. Report dosage or exposure, such as minutes delivered or learning opportunities presented, in its own field. A session can have the planned duration while several components occur incorrectly.
For IOA, two trained observers watch the same defined period and record independently. Predefine the observation window, event-matching rule, tolerance for timing or duration, calculation, and treatment of missing records. If a human observer scores treatment integrity, sample IOA on those integrity records too. Agreement on client behavior leaves the fidelity measure unexamined.
The current BACB Ethics Code page identifies the governing code for BCBA and BCaBA certificants and applicants. The Ethics Code for Behavior Analysts addresses appropriate data procedures, continual evaluation, conditions that interfere with service delivery, stakeholder involvement, and performance monitoring. The public CASP ABA Practice Guidelines summary describes Version 3.0 as guidance for planning, implementing, and evaluating ABA assessment and treatment for autism. Access to the full guideline requires a CASP licensing agreement. The public summary supplies no single integrity or IOA threshold and should not be cited as one.
Use integrity measures that expose important errors
A global percentage is useful for orientation. Component-level results show where implementation diverged. The distinction matters most when components differ in clinical consequence.
In an empirical report, four undergraduate therapists each implemented the same 12-trial discrete-trial teaching program with one child in one school setting. Cook and colleagues found that global integrity scores masked low performance on individual components. For Phillip, global integrity averaged 86.5% across the last three feedback sessions, while integrity for securing attention, recording data, delivering reinforcers, and implementing error correction was 29.9%, 37.5%, 37.5%, and 58.3%, respectively. This small study supports examining components alongside a global summary.
Build the integrity record around these layers:
- Global accuracy: all correct component opportunities divided by all applicable opportunities.
- Component accuracy: correct opportunities for one component divided by opportunities for that component.
- Critical-step status: each safety, medical, assent, dissent, AAC-access, or restrictive-procedure safeguard recorded separately.
- Error type: omissions, additions, incorrect timing, incorrect sequence, excessive intensity, or unauthorized adaptation.
- Dose and context: exposure, implementer, setting, phase, client signals, staffing, materials, competing demands, and relevant environmental events.
Set opportunity rules before scoring. An opportunity that ends because the clinician correctly honors assent withdrawal can be coded as withdrawal and removed from the remaining teaching-step denominator. Honoring the withdrawal may itself be a scored safeguard. Missing materials count as an implementation error when preparing or making them available belongs to the implementer's defined responsibility. Every exclusion should remain visible.
Foreman and colleagues first observed teachers' omission and commission errors during natural implementation of existing timeout-from-play procedures for four children. In a second study, a researcher used a reversal design with two children to compare no timeout, 100% integrity, and the omission-integrity levels observed in the first study. Reduced-integrity timeout decreased the measured behavior relative to baseline for both children, within the study's sequence and conditions. That result illustrates why a percentage alone cannot predict a client's outcome or identify which component matters. Treatment selection and restrictive procedures require their own clinical, consent, risk, and review processes.
Match the IOA calculation to the data
The calculation should be sensitive to the disagreement most capable of changing the decision.
Data and riskUseful calculationCalculation and cautionTotal event countTotal-count IOASmaller count divided by larger count, multiplied by 100. Identical totals can hide different event timing. Two zero counts describe no recorded events; label the denominator rather than converting it automatically to 100%.Count across intervalsMean count-per-interval IOAWithin each interval, divide smaller by larger and average the interval percentages. State how zero-zero intervals are handled.Exact count across intervalsExact count-per-interval IOAIntervals with identical counts divided by all compared intervals. This is stricter than comparing session totals.Occurrence or nonoccurrence by intervalInterval-by-interval IOAIntervals with the same occurrence status divided by compared intervals. Many shared empty or occupied intervals can inflate agreement.Low-rate occurrenceScored-interval IOAAgreements on occurrence divided by intervals in which either observer recorded occurrence. This focuses the result on detected events.High-rate occurrenceUnscored-interval IOAAgreements on nonoccurrence divided by intervals in which either observer recorded nonoccurrence. This focuses the result on detected pauses.Trial outcome or categoryTrial-by-trial IOATrials with the same recorded result divided by compared trials. Define how observers match trials and multi-response categories.DurationTotal-duration IOA or mean duration-per-occurrence IOATotal-duration IOA is the shorter total duration divided by the longer total duration, multiplied by 100; it can hide event-level disagreement. For mean duration-per-occurrence IOA, match corresponding occurrences using the prespecified rule, calculate shorter duration divided by longer duration for each matched occurrence, and average those percentages. State how unmatched events and two zero totals are handled.
The BACB BCBA Test Content Outline, Sixth Edition is an examination blueprint. It identifies evaluating measurement validity and reliability, selecting representative procedural-integrity measures that account for accuracy and dosage, and making data-based decisions about integrity and intervention effectiveness as separate entry-level BCBA tasks. It creates no clinical formula, sampling percentage, or pass score.
Vollmer, Sloman, and St. Peter Pipkin's practice article on data reliability and treatment-integrity monitoring explains how both measures affect clinicians' ability to judge intervention effects and offers practical calculations. Apply each calculation with the assumptions and limitations visible. IOA is agreement. Accuracy requires a valid reference, sound definitions, representative sampling, and other evidence suited to the decision.
Plan representative samples before seeing the scores
Write the sampling plan before selecting sessions. Spread integrity and IOA observations across relevant:
- baseline, teaching, maintenance, generalization, and transition phases
- implementers, supervisors, settings, times, and service formats
- common and difficult procedures
- high-risk and critical components
- low-rate and high-rate target conditions
- sessions with known staffing, material, health, or environmental barriers
Select routine observations from a defined eligible pool using random selection or a prespecified stratified schedule applied reproducibly. This reduces convenience sampling; reproducibility alone does not make a sample representative. Triggered observations remain useful after an incident, complaint, unusual graph, plan change, new implementer, or suspected drift. Report triggered and routine samples separately.
Choose frequency and action rules from the consequence of an error, variability in performance, observer experience, procedure complexity, client risk, data volume, and the decision the measure will support. A single favorable probe supplies weak evidence for a treatment-period conclusion. A practice-wide percentage can also hide one implementer, setting, phase, or component with recurring problems.
Keep observers independent during scored observations. They can calibrate on separate examples beforehand. During the scored sample, each observer uses their own record and receives no coaching or access to the other's data. Label a coached observation as training. When video is used, obtain required consent or authorization, use an approved secure system, limit recording and access, and follow retention and deletion rules.
A 2023 review by Essig, Rotta, and Poling examined experimental articles in two behavior-analytic journals from 2017 through 2021. The authors reported procedural-fidelity data in 54.7% of relevant articles. Of articles reporting fidelity for a human-implemented independent variable, 17.7% also reported IOA for fidelity. Separately, among articles whose dependent variable was measured by a human observer, 96.4% reported IOA for participant behavior. These publication findings support checking human-scored fidelity and establish no clinical sampling quota.
Interpret treatment integrity and IOA together
PatternWhat the clinician can sayImmediate reviewHigher integrity, higher IOAThe measured procedure was delivered with the reported fidelity, and observers showed the reported agreement in sampled periods.Examine client outcomes, unwanted effects, assent, social validity, contextual fit, representativeness, and alternative explanations.Higher integrity, lower IOAImplementation may be strong, while the observation record lacks enough consistency for the pending decision.Repair definitions, event matching, observer training, independence, calculation, and sampling; collect a fresh probe.Lower integrity, higher IOAObservers consistently detected a delivery problem.Examine component errors, feasibility, materials, workload, training, plan clarity, environmental support, and client fit.Lower integrity, lower IOAThe team has uncertain implementation and uncertain measurement.Stabilize immediate safety and communication safeguards, then repair both systems before attributing the outcome.
High results on both measures support a narrower conclusion than “the treatment works.” A plan can be implemented precisely and measured consistently while producing little benefit, undue burden, or unwanted effects. Continual evaluation still includes the client's progress and experience, caregiver or stakeholder input, health variables, assent and dissent, AAC access, generalization, maintenance, and treatment alternatives.
A fictional example shows why the method matters
A fictional supervisor observes 10 teaching trials with five applicable components per trial. That creates 50 component opportunities. The implementer completes 47 correctly:
47 ÷ 50 × 100 = 94% global treatment integrity
One of the three errors is a missed critical AAC-access and assent check. The audit reports 94% globally, the five component results, and the critical miss. The team follows its critical-step route immediately. The average never clears that safeguard.
Two independent observers also record whether the client initiates a help request on each trial. Observer A records occurrence on trials 1, 2, 4, 6, 8, and 10. Observer B records occurrence on trials 1, 3, 4, 6, 8, and 10. Both record six events, so total-count IOA is:
6 ÷ 6 × 100 = 100%
They agree on eight of the 10 trial outcomes, so trial-by-trial IOA is:
8 ÷ 10 × 100 = 80%
The 80% result is a descriptive score, not a pass line. The team checks the definition and video timestamps, finds that the two observers applied the response-window boundary differently, clarifies that boundary, recalibrates on a new fictional example, and collects a fresh independent probe. It preserves the original data and audit trail.
Use one decision sequence every time
- Confirm the decision. State whether the data will guide a routine adjustment, high-risk procedure, caregiver training, staff competency decision, authorization request, or another action.
- Check measurement confidence. Inspect definitions, observer independence, sampling, calculation, raw records, missing data, and IOA for the relevant measure.
- Check implementation. Review global, component, critical-step, error-type, dosage, implementer, phase, and context data.
- Check the plan and outcome. Examine assessment logic, client priorities, assent or dissent, communication access, progress, unwanted effects, health and environmental variables, feasibility, and stakeholder input.
- Act and recheck. Assign the correction, responsible person, deadline, next independent sample, clinical review date, and communication or consent step.
When IOA is weak for a high-consequence decision, preserve the disagreement and gather stronger evidence before relying on the disputed measure. When integrity is weak, investigate why. A clearer procedure, better materials, competency-based training, workload change, environmental support, or clinically appropriate plan revision may fit the finding. Repeated demands for perfect execution can consume time while leaving an unworkable plan in place.
Treatment integrity and IOA audit checklist
Before using either score, confirm that the record shows:
- [ ] the active plan version and observation date
- [ ] observable components and opportunity definitions
- [ ] permitted adaptations and error categories
- [ ] component, global, critical-step, and dosage results where relevant
- [ ] client assent, dissent, AAC access, safety, and contextual events
- [ ] observer names or coded identifiers, training, and independence
- [ ] matched observation windows and raw records
- [ ] the IOA calculation, denominator, zero-event rule, and event-matching tolerance
- [ ] routine versus triggered sampling and distribution across relevant conditions
- [ ] exclusions, missing data, invalid opportunities, and recording limitations
- [ ] the client's outcome data and experience considered alongside the two scores
- [ ] a source-linked action rule, owner, due date, and recheck
- [ ] required privacy, recording, documentation, supervision, payer, and state controls
Treat integrity observations as clinical quality data, not a character score. Feedback should identify the observable component, context, evidence, consequence, support, and next probe. The credentialed clinician responsible for the case retains accountability for treatment design, risk decisions, data interpretation, and any clinical change within applicable scope and supervision rules.
Related resources
- Data, Outcomes and Clinical Decision-Making
- When an ABA Client Is Not Making Progress: A Structured Clinical Review
- How to Read ABA Graphs and Make Defensible Treatment Decisions
- Concurrent ABA Billing: A Decision Guide for Overlapping Services
- How to Build Effective ABA Caregiver Training With Behavioral Skills Training
Sources
- Behavior Analyst Certification Board, Ethics Codes
- Council of Autism Service Providers, ABA Practice Guidelines Version 3.0 public summary and licensing information
- Behavior Analyst Certification Board, Ethics Code for Behavior Analysts
- Behavior Analyst Certification Board, BCBA Test Content Outline, Sixth Edition
- Vollmer, Sloman, and St. Peter Pipkin, Practical Implications of Data Reliability and Treatment Integrity Monitoring
- Essig, Rotta, and Poling, Interobserver Agreement and Procedural Fidelity: An Odd Asymmetry
- Cook and colleagues, Global Measures of Treatment Integrity May Mask Important Errors in Discrete-Trial Training
- Foreman and colleagues, Treatment Integrity Failures During Timeout From Play
- Reed and Azulay, A Microsoft Excel 2010 Based Tool for Calculating Interobserver Agreement
- Centers for Disease Control and Prevention, CASPER Sampling Methodology