Reliability is the consistency of your measurement. In single case design the central question is whether a second, independent observer would score your sessions the same way you did. That check is interobserver agreement. Expand each panel for who does it, how much, how it is calculated, and the standard you must meet.
Use the Single-Case Design Rigor Planner to plan your fidelity, reliability, and validity decisions for the proposal.
1 Interobserver agreement (IOA) ▼

Interobserver agreement is the degree to which two observers, scoring the same sessions independently, produce the same data. It is the form reliability takes in single case design. High agreement tells a reader that your data reflect the behavior, not one scorer's habits.

Reliability is necessary but not sufficient for validity. A measure can be perfectly consistent and still measure the wrong thing. You need reliability to trust your data, and you need validity for that data to mean what you claim. Getting agreement high does not, by itself, make the study valid.

Other reliability ideas from measurement in general, such as test-retest reliability, internal consistency, and parallel forms, describe consistency of a test instrument over time or across items. In an intervention study built on repeated direct observation, interobserver agreement is the reliability you plan and report.
2 Who the second observer is ▼

The second observer scores a portion of your sessions independently, without seeing your scores. Choose and prepare that person deliberately.

  • Independent: The second observer scores on their own and does not compare with you until both are done.
  • Trained to a criterion: Train the observer on your operational definitions and data sheet, then practice until agreement reaches a set level before any study data are scored.
  • Blind where possible: An observer who does not know the phase or condition of a session cannot be swayed by expecting a change.
Example of training to criterion In the worked narrative study, training ran in three stages: the second coder watched while the researcher transcribed, then both transcribed the same session independently and compared, then repeated that independent-and-compare step until discrepancies were resolved. Coders practiced until they reached the agreement standard the study required before scoring real data.
Example of blinding Namasivayam and colleagues (2024) measured intelligibility with naive listeners who were blind to participant details, disorder, and whether a recording came from before or after treatment, so that expectation could not shape the scores.
3 How much to double score ▼

You do not need a second observer on every session. The default in single case design is to have a second observer independently score at least 20% of sessions. More is stronger.

  • Spread the double-scored sessions across every phase and condition, not just baseline or just intervention, so agreement is demonstrated everywhere you draw conclusions.
  • Select the double-scored sessions at random rather than picking your cleanest ones.
  • Report the exact percentage of sessions that were double scored.
Example Namasivayam and colleagues (2024) had a licensed speech-language pathologist independently re-transcribe about 40% of the entire data set for reliability, well above the 20% minimum.
4 Calculation methods ▼

Report which method you used and the value you obtained. The right method depends on what you are counting.

MethodWhat it doesFits
Point-by-point agreement Compares the two observers' scores instance by instance, then reports the percentage that agreed Coded behaviors where each instance can be lined up, for example scoring each utterance or each trial
Total count (smaller over larger) Divides the smaller total count by the larger total count for a session A quick session-level check on frequency counts, weaker than point-by-point because it can hide offsetting errors
Cohen's kappa Agreement on categorical judgments corrected for the agreement expected by chance Yes or no or category coding where chance agreement would otherwise inflate the number
Intraclass correlation (ICC) Agreement among raters on continuous or scaled scores Measures that produce a number or rating rather than a category

Point-by-point agreement is the workhorse for coded single case data. It is calculated as:

agreement = agreements / (agreements + disagreements) x 100

To compute your chosen agreement value, open the Single Case Inferential Statistics Guide and choose the Reliability option, which walks you through the agreement statistic that fits your data.

Worked example Namasivayam and colleagues (2024) reported reliability of broad transcription using a point-by-point agreement index, obtaining an average of 86.4% across the roughly 40% of the data set that was re-scored. That value is above the 80% standard, so the transcription data are considered reliable.
Watch out Kappa and ICC correct for chance and for scale, so a raw percentage agreement and a kappa are not the same number and are not interchangeable. State which one you report.
5 The 80% standard and corrective action ▼

The default minimum for acceptable interobserver agreement is 80%. If agreement falls below that, you do not simply report a low number and move on. You take corrective action and document it.

  • Hold a discrepancy discussion: the observers review the sessions where they disagreed and identify why.
  • Sharpen the operational definition or the data sheet if the disagreement traces to an unclear rule.
  • Retrain and re-establish agreement on practice sessions before scoring more study data.
Example of a corrective process In the worked narrative study, when the first attempt did not reach the agreement standard, the coders used a two-step consensus process: one coder scored the full measurement portion, then both reviewed the session together and resolved every disagreement, with a set rule for who decided in the rare case of a standoff. They report reaching full agreement across sessions once the process was in place. The corrective step, not the first imperfect pass, is what makes the data defensible.
6 What to write in your Methods chapter ▼

Your reliability section should let a reader judge how trustworthy your data are. Include each of these:

  • Who the second observer was and how they were trained, including the training criterion.
  • The percentage of sessions double scored, and confirmation that they spanned phases and were selected at random.
  • The calculation method and the agreement value obtained.
  • The minimum you set (80% or higher) and what you did if a session fell below it.

The SCRIBE reporting guideline, included in this course, lists the reliability details reviewers expect a single case report to state.

□ Sources ▼