Showing that your measurement is consistent, not just that it happened.
Interobserver agreement is the degree to which two observers, scoring the same sessions independently, produce the same data. It is the form reliability takes in single case design. High agreement tells a reader that your data reflect the behavior, not one scorer's habits.
Reliability is necessary but not sufficient for validity. A measure can be perfectly consistent and still measure the wrong thing. You need reliability to trust your data, and you need validity for that data to mean what you claim. Getting agreement high does not, by itself, make the study valid.
The second observer scores a portion of your sessions independently, without seeing your scores. Choose and prepare that person deliberately.
You do not need a second observer on every session. The default in single case design is to have a second observer independently score at least 20% of sessions. More is stronger.
Report which method you used and the value you obtained. The right method depends on what you are counting.
| Method | What it does | Fits |
|---|---|---|
| Point-by-point agreement | Compares the two observers' scores instance by instance, then reports the percentage that agreed | Coded behaviors where each instance can be lined up, for example scoring each utterance or each trial |
| Total count (smaller over larger) | Divides the smaller total count by the larger total count for a session | A quick session-level check on frequency counts, weaker than point-by-point because it can hide offsetting errors |
| Cohen's kappa | Agreement on categorical judgments corrected for the agreement expected by chance | Yes or no or category coding where chance agreement would otherwise inflate the number |
| Intraclass correlation (ICC) | Agreement among raters on continuous or scaled scores | Measures that produce a number or rating rather than a category |
Point-by-point agreement is the workhorse for coded single case data. It is calculated as:
To compute your chosen agreement value, open the Single Case Inferential Statistics Guide and choose the Reliability option, which walks you through the agreement statistic that fits your data.
The default minimum for acceptable interobserver agreement is 80%. If agreement falls below that, you do not simply report a low number and move on. You take corrective action and document it.
Your reliability section should let a reader judge how trustworthy your data are. Include each of these:
The SCRIBE reporting guideline, included in this course, lists the reliability details reviewers expect a single case report to state.