Single-Case Design: Rigor Planner

Work through the fidelity, reliability, and validity decisions for your single-case study before you write your proposal. It makes sure you have made every decision and can defend it. When you finish, generate a summary of your decisions and use it to draft each section of your paper.

Saving and resuming

This is a long set-up, and your work is not stored automatically. To stop and continue later, click Save progress at the bottom of the page. That downloads a file named rigor-planner-progress.json. When you come back, open this page again, click Load saved progress, and choose that file to pick up exactly where you left off. Keep the file somewhere you will find it, and save again whenever you make changes you want to keep.

1Treatment Fidelity

Fidelity documents that the intervention (your independent variable) was actually delivered the way you designed it. Fidelity is also part of reliability: it is the reliability of how consistently the independent variable was implemented. You may see it called treatment fidelity, procedural integrity, or treatment integrity. These terms refer to the same idea.

Name the specific instrument that lists the intervention's essential components. If you are adapting or creating one, say so.

Consider bias: the closer the observer is to the intervention, the more independent verification your reviewers will expect.

You must monitor at least 20% of sessions, but you cannot observe a fraction of a session, so round up to the next whole session. With a small number of sessions that pushes the actual percentage above 20%, and you report the real percentage for each phase separately. Enter your session counts and the tool does the math.

Phase or conditionTotal sessions20% (raw)Monitor (rounded up)Actual %
Research norm: fidelity is usually monitored on 20% or more of sessions, reported separately for each phase.

Select the approach you will use. Click a term to see what it means and a communication-disorders example.

Percent adherence

The proportion of the intended intervention steps or components that were actually delivered. It counts whether each component happened, not how well.

Example: An aided language modeling protocol has 8 steps; the clinician completes 7 in a session, so adherence is 88%.
Quality-weighted adherence

Adherence adjusted for how well each component was delivered, not just whether it occurred. Each step is rated on a quality scale, so a component delivered poorly earns partial credit.

Example: Two clinicians each deliver a recast during a language session, but one is contingent and well-timed while the other is delayed and off-target; quality weighting gives them different credit.
Dosage fidelity

Whether the intended amount of intervention was delivered: the number of sessions, their length, or the frequency of a target procedure.

Example: A phonology protocol calls for 100 production trials per session; a session with only 60 trials has high procedural adherence but low dosage fidelity.
Composite score

A single index that combines two or more dimensions, such as adherence, quality, and dosage, into one fidelity figure. Define how the pieces are weighted and combined so the number is interpretable.

Example: A fluency intervention combines procedural adherence, the quality of the clinician's feedback, and the number of practice trials into one fidelity index.
Research norm: the minimum is usually 80% or higher, though it may be set higher or lower depending on the measure and what a component's failure would mean for the intervention.

Baseline and intervention are both monitored

Fidelity applies to more than the intervention phase. You need a fidelity check for baseline (that baseline conditions were held as planned, with the intervention withheld) and a fidelity check for intervention. Describe each below.

Sources you can cite for this section

Cite a source for your fidelity tool and approach. Options from your course readings:

Anis, L., Benzies, K. M., Ewashen, C., Hart, M. J., & Letourneau, N. (2021). Fidelity assessment checklist development for community nursing research in early childhood. Frontiers in Public Health, 9, 582950. https://doi.org/10.3389/fpubh.2021.582950

Wolery, M., Dunlap, G., & Ledford, J. R. (2011). Single-case experimental methods: Suggestions for reporting. Journal of Early Intervention, 33(2), 103–109. https://doi.org/10.1177/1053815111418235

Tate, R. L., & Perdices, M. (2020). Research note: Single-case experimental designs. Journal of Physiotherapy, 66, 202–206. https://doi.org/10.1016/j.jphys.2020.06.004

Kratochwill, T. R., Hitchcock, J., Horner, R. H., Levin, J. R., Odom, S. L., Rindskopf, D. M., & Shadish, W. R. (2010). Single-case designs technical documentation. What Works Clearinghouse. https://ies.ed.gov/ncee/wwc/pdf/wwc_scd.pdf

2Reliability

Reliability establishes that your measurement of the dependent variable is consistent and not an artifact of a single observer. You will plan who double-scores, how much gets double-scored, and how agreement is calculated. Remember that treatment fidelity (Section 1) is also a form of reliability, applied to the independent variable.

Select the type or types your design uses. Click a term for its definition and an example.

Interobserver agreement (IOA)

The degree to which two independent observers, watching the same behavior at the same time, record the same thing. It is the most common reliability index in single-case behavioral measurement and is what the What Works Clearinghouse standards refer to.

Example: Two observers independently score the same video of a child's requesting during a milieu teaching session and compare how often each recorded a request.
Interrater reliability

The consistency of scores when two or more raters apply a rating scale or coding system to the same performance. Often used when the measure involves judgment along a scale rather than a simple count.

Example: Two SLPs rate the same narrative sample on a 0-4 story-grammar rubric and their ratings are compared.
Interassessor agreement

Agreement between two people who independently administer or score an assessment, such as a standardized probe or criterion measure. Relevant when the dependent variable comes from an assessment rather than live observation.

Example: Two examiners independently administer and score the same standardized articulation probe and their scores are compared.

State their training and, where possible, whether they are blind to phase or condition.

Same math as fidelity: at least 20%, rounded up to whole sessions, reported for each phase and condition. WWC requires reliability in every phase, so add a row for each phase your design has.

Phase or conditionTotal sessions20% (raw)Double-score (rounded up)Actual %
Research norm: reliability is typically collected on at least 20% of scores in every phase and condition, not clustered in one.

This determines which agreement methods are appropriate. Pick the one that matches how you record the dependent variable.

Select your recording system above to narrow these to the methods that fit your data. Definitions for every method stay available below.

Total count / total agreement

Divide the smaller total count by the larger total count for a session and multiply by 100. It compares only the two overall totals, so it is the most lenient method.

Example: Observer A counts 20 spontaneous words and Observer B counts 18; total agreement is 18/20 = 90%, even if they did not agree on which specific words.
Point-by-point (exact, trial-by-trial) agreement

Compare the two observers trial by trial and count a trial as an agreement only when both recorded the same thing on that trial. Agreements over agreements plus disagreements, times 100. Far stricter than total count.

Example: On a discrete-trial correct/incorrect record of /s/ productions, only trials where both observers scored the same production count as agreements.
Interval-by-interval agreement

For interval recording, compare the two records interval by interval and count every interval, scored or unscored, on which they agree. Can inflate agreement when the behavior is very rare or very frequent.

Example: During 10-second partial-interval recording of stuttering, the two records are compared across all intervals.
Scored-interval agreement

Count agreement only among intervals where at least one observer scored the behavior as occurring. Controls for inflation when the behavior is infrequent.

Example: For a low-rate behavior like an initiation of joint attention, agreement is counted only in intervals where at least one observer scored an initiation.
Unscored-interval agreement

Count agreement only among intervals where at least one observer scored the behavior as not occurring. Controls for inflation when the behavior is very frequent.

Example: For a high-rate behavior like on-task looking during a session, agreement is counted only in intervals where at least one observer scored it as not occurring.
Mean count-per-interval agreement

Compute agreement on the count within each interval, then average those interval agreements across the session. More sensitive than total count.

Example: Vocalizations are counted within each 1-minute interval, agreement is computed for each interval, then averaged across the session.
Cohen's kappa

Adjusts percentage agreement for the agreement that would be expected by chance alone. It is considered the more psychometrically sound option and is used for categorical or nominal coding systems, so it is typically lower than raw percentage agreement for the same data.

Example: When two observers code each utterance as a request, comment, or protest, kappa corrects for the agreement expected if they had simply guessed categories at the base rates.
Correlational reliability (Pearson r and/or ICC)

Used for continuous, session-level scores rather than trial or interval agreement, for example comparing two observers' overall session scores across a series of sessions. Pearson r indexes how well the scores track together; the intraclass correlation coefficient (ICC) also accounts for agreement in absolute value, not just rank order.

Example: Two observers each assign an overall intelligibility percentage to 15 speech samples; Pearson r or ICC indexes how closely their session-level scores track across the 15 samples.

Tie your choice to your measurement system: count versus interval versus continuous data, how frequent the behavior is, and whether coding is categorical. Cite a source for your reasoning.

References you can cite for your rationale

Kratochwill, T. R., Hitchcock, J., Horner, R. H., Levin, J. R., Odom, S. L., Rindskopf, D. M., & Shadish, W. R. (2010). Single-case designs technical documentation. What Works Clearinghouse. https://ies.ed.gov/ncee/wwc/pdf/wwc_scd.pdf

Wolery, M., Dunlap, G., & Ledford, J. R. (2011). Single-case experimental methods: Suggestions for reporting. Journal of Early Intervention, 33(2), 103–109. https://doi.org/10.1177/1053815111418235

Viera, A. J., & Garrett, J. M. (2005). Understanding interobserver agreement: The kappa statistic. Family Medicine, 37(5), 360–363.

Walter, S. R., Dunsmuir, W. T. M., & Westbrook, J. I. (2019). Inter-observer agreement and reliability assessment for observational studies of clinical work. Journal of Biomedical Informatics, 100, 103317. https://doi.org/10.1016/j.jbi.2019.103317

Vanbelle, S., Hernandez Engelhart, C., & Blix, E. (2025). Measuring agreement in diagnostics: A practical guide for researchers. Statistics in Medicine. Advance online publication.

What Works Clearinghouse: single-case design standards

The What Works Clearinghouse single-case design standards require interobserver agreement to be collected on at least 20% of data points in each phase and condition, with minimum acceptable values of at least 80% for percentage agreement, or at least 0.60 for kappa. Plan your reliability so it can meet this standard in every phase, not just overall.

3Validity

Validity is not something you bolt on at the end. As you design your study, you build it in. Work through the four kinds of validity below and decide how your design handles each. Anywhere your design is weak, note it now: you will address it as a limitation when you present, explaining how and why it applies.

Internal validity

Internal validity asks whether your intervention, and not some other factor, caused the change you observed. In single-case design this is established mainly through replicating the effect, rather than through comparison groups. Decide which of the following your design uses.

Active manipulation of the independent variable

You, the researcher, decide when to start or change the intervention. Because you control the timing, you can be sure the intervention came before the change you measured, which is what lets you argue it caused the change.

Example: The researcher decides the exact session at which a child's AAC device is introduced, rather than letting it appear whenever a parent happens to bring it.
Replication of effect

The design shows the same effect at least three separate times, at three different points, whether within one case or across cases. Seeing the effect repeat when, and only when, the intervention is introduced is the core of the single-case argument for causation.

Example: In a multiple baseline across three children learning a core-vocabulary board, the effect is demonstrated three times as the intervention is staggered to each child.
Baseline logic / steady-state strategy

Establish a stable baseline, typically 3 to 5 data points, before you intervene. When the baseline is stable and change appears only after intervention, that change is more attributable to the treatment than to a pre-existing trend or to ordinary variability.

Example: Collecting 3 to 5 stable baseline probes of correct /r/ production before treatment begins, so a later rise is not just a pre-existing upward trend.
Randomization

Randomization controls threats that replication alone does not fully address, and it supports randomization-based statistical tests. Examples include randomizing the order of phases, the start points, or the assignment of conditions.

Example: Randomly assigning which of three target phonemes enters treatment first in a multiple baseline across behaviors.
A priori decision rules

Criteria you specify in advance, before data collection begins, for when phases change, which participants are included or excluded, and when a baseline counts as stable. Setting these before you see the data reduces confirmation bias, a real risk in response-guided designs.

Example: Specifying in advance that a phase changes only after three consecutive stable data points, decided before any child's data are seen.

Click each threat to see what it means, a communication-disorders example, and a box to write what your design does about it. Fill in the ones your design needs to handle. Only the threats you write about will appear in your summary.

History

An outside event, other than the intervention, happens during the study and could explain the change.

Example: A child also starts a new classroom reading program mid-study; that event, not the intervention, might explain gains. Staggered baselines protect against this, because one classroom change cannot explain gains that appear in each child only when treatment reaches them.
Maturation

Natural changes within the participant over time, such as development, learning, fatigue, or growth, that could account for the change instead of the intervention.

Example: A toddler's mean length of utterance may rise through normal language development over the study, independent of the intervention.
Testing

The act of being measured repeatedly changes performance, so improvement reflects practice with the measure rather than the intervention.

Example: Repeatedly probing the same vocabulary words could raise scores through practice with the probe itself.
Instrumentation

The measurement tool or the observers themselves drift over time, so a change reflects a shift in how you measured rather than a real change in behavior.

Example: An observer becomes more lenient about scoring words as intelligible as the study goes on, so the measure drifts rather than the behavior.
Statistical regression

Also called regression to the mean. When participants are selected because of extreme scores, those scores tend to move toward the average on re-measurement, which can look like an effect but is not one.

Example: A child selected because of an unusually low articulation probe on one bad day may score higher next time simply through regression to the mean.
Attrition

Participants drop out during the study in a way that biases the results, for example if those who leave differ systematically from those who stay.

Example: If the two children who make the least progress drop out of an AAC study, the remaining data overstate the intervention's effect.
Selection

Pre-existing differences between who ends up in which condition, so a difference in outcomes reflects who the participants were rather than what the intervention did.

Example: Assigning more verbal children to one condition and less verbal children to another would confound who they were with what the intervention did.
Ambiguous temporal precedence

It is unclear which variable came first, so you cannot tell whether the intervention caused the change or the reverse.

Example: If a parent began home language stimulation at the same time treatment started, it is unclear which came first; a clear baseline with researcher-controlled onset resolves it.
Sources you can cite for internal validity

Cite a source when you claim your design controls a threat. Options from your course readings:

Slocum, T. A., Joslyn, P. R., Nichols, B., & Pinkelman, S. E. (2022). Revisiting an analysis of threats to internal validity in multiple baseline designs. Perspectives on Behavior Science, 45, 681–694. https://doi.org/10.1007/s40614-022-00351-0

Kratochwill, T. R., Hitchcock, J., Horner, R. H., Levin, J. R., Odom, S. L., Rindskopf, D. M., & Shadish, W. R. (2010). Single-case designs technical documentation. What Works Clearinghouse. https://ies.ed.gov/ncee/wwc/pdf/wwc_scd.pdf

Tate, R. L., & Perdices, M. (2020). Research note: Single-case experimental designs. Journal of Physiotherapy, 66, 202–206. https://doi.org/10.1016/j.jphys.2020.06.004

Smith, J. D. (2012). Single-case experimental designs: A systematic review of published research and current standards. Psychological Methods, 17(4), 510–550. https://doi.org/10.1037/a0029312

External validity

External validity, the generalization of findings, is a challenge in single-case design because the number of participants is small. Full direct and systematic replication happens across separate studies and is beyond the scope of this proposal, but you can build several achievable routes to generality into your own study. Select the ones you will use.

Within-study replication

A multiple baseline across several participants, behaviors, or settings demonstrates the effect more than once under the same procedures. This is direct-replication logic contained inside a single study, which is what makes it achievable in your proposal.

Example: A multiple baseline across four children shows the effect four times as the same script-fading procedure reaches each child in turn.
Detailed participant and setting description

Reporting participant and setting characteristics in enough detail that a reader can judge to whom, and to what contexts, the results might extend. In small-N work this description is a real contributor to external validity, not just background.

Example: Reporting each child's age, diagnosis, language level, and the clinic or classroom setting so readers can judge who the findings might apply to.
Direct replication (across studies – informational)

Repeating the same study with no procedural changes, to confirm the finding is reliable and reproducible. This happens across separate studies, so it is background for how a literature builds generality rather than something you carry out in this proposal.

Example: Another clinician repeats the same script-fading study with a new child with autism to confirm the finding holds.
Systematic replication (across studies – informational)

Deliberately varying participants, settings, behaviors, or interventionists to test the boundaries of generalization. Also an across-studies activity, included here so you understand where your single study sits in the larger evidence base.

Example: Repeating a milieu teaching study with preschoolers instead of toddlers, or in a classroom instead of a clinic, to test how far the effect generalizes.

Generality is not only about N

A small number of participants does not by itself doom external validity. Within-study replication, planned generalization, and careful description all build generality from a single study. See Walker & Carr (2021), Generality of Findings From Single-Case Designs: It's Not All About the N, in your course readings.

Sources you can cite for external validity

Cite a source to support your generality argument. Options from your course readings:

Walker, S. G., & Carr, J. E. (2021). Generality of findings from single-case designs: It's not all about the N. Behavior Analysis in Practice, 14, 991–995. https://doi.org/10.1007/s40617-020-00547-3

Machalicek, W., Gross, D. P., Armijo-Olivo, S., Ferriero, G., Kiekens, C., Martin, R., Walshe, M., & Negrini, S. (2024). The role of single case experimental designs in evidence creation in rehabilitation. European Journal of Physical and Rehabilitation Medicine, 60(6), 1100–1111. https://doi.org/10.23736/S1973-9087.24.08713-6

Dallery, J., & Raiff, B. R. (2014). Optimizing behavioral health interventions with single-case designs: From development to dissemination. Translational Behavioral Medicine, 4(3), 290–303. https://doi.org/10.1007/s13142-014-0258-z

Social validity

Social validity asks whether the goals, the procedures, and the outcomes of the intervention are acceptable and meaningful to the people affected by them: the client, the family, teachers, or clinicians. It is separate from whether the intervention statistically works. A treatment can produce a reliable effect that stakeholders still find impractical or unimportant. It is assessed through consumer or stakeholder ratings, satisfaction measures, or comparison to socially important benchmarks, such as whether a child's outcome now matches that of typically developing peers. If you address it, add a separate social-validity section toward the end of your methods, with its own measure.

See a communication-disorders example
Example: Asking parents and the classroom teacher to rate whether a child's new AAC-based requesting is understandable and useful in daily routines, and whether the gains matter to them.
Sources you can cite for social validity

If you address social validity, cite a source for the construct and your measure. Options from your course readings:

Wolery, M., Dunlap, G., & Ledford, J. R. (2011). Single-case experimental methods: Suggestions for reporting. Journal of Early Intervention, 33(2), 103–109. https://doi.org/10.1177/1053815111418235

Tate, R. L., & Perdices, M. (2020). Research note: Single-case experimental designs. Journal of Physiotherapy, 66, 202–206. https://doi.org/10.1016/j.jphys.2020.06.004

Dallery, J., & Raiff, B. R. (2014). Optimizing behavioral health interventions with single-case designs: From development to dissemination. Translational Behavioral Medicine, 4(3), 290–303. https://doi.org/10.1007/s13142-014-0258-z

Statistical conclusion validity

In single-case design, visual analysis of the graphed data is the dominant method, supplemented by effect-size statistics. For this assignment, visual analysis and effect-size statistics are both required. Randomization tests are an optional addition.

Visual analysis: the five features

Level is the average value of the data within a phase. Trend is the slope or direction within a phase. Variability is how much the data bounce around their level. Immediacy of effect is how quickly the data change when a phase changes; a fast change is stronger evidence. Overlap is how many data points in adjacent phases share the same values; less overlap means a stronger effect.

Example: Judging a child's /k/ production graph by its level, trend, variability, how immediately it changed at the phase line, and how much baseline and treatment points overlap.
Effect-size statistics

Quantitative summaries of the magnitude of change that supplement visual analysis, helping guard against serial dependence and autocorrelation.

Example: Reporting Tau-U or a between-case standardized mean difference to quantify the size of the change in correct productions.
Randomization tests

Tests that use the random elements of your design, such as randomly assigned start points, to test whether an observed effect is larger than chance would produce.

Example: Using a randomization test built on randomly assigned intervention start points to test whether the observed change exceeds chance.
Serial dependence & autocorrelation

Successive data points in a time series are often correlated with each other rather than independent. This is serial dependence, and autocorrelation is its measure. It violates the independence assumption behind many statistics, which is why single-case analysis leans on visual analysis and on tests built to handle it.

Example: Consecutive daily probes of the same child tend to be correlated, so ordinary tests that assume independent data points can mislead.
Sources you can cite for statistical conclusion validity

Cite a source for your analysis approach. Options from your course readings:

Kratochwill, T. R., Hitchcock, J., Horner, R. H., Levin, J. R., Odom, S. L., Rindskopf, D. M., & Shadish, W. R. (2010). Single-case designs technical documentation. What Works Clearinghouse. https://ies.ed.gov/ncee/wwc/pdf/wwc_scd.pdf

Smith, J. D. (2012). Single-case experimental designs: A systematic review of published research and current standards. Psychological Methods, 17(4), 510–550. https://doi.org/10.1037/a0029312

Wolery, M., Dunlap, G., & Ledford, J. R. (2011). Single-case experimental methods: Suggestions for reporting. Journal of Early Intervention, 33(2), 103–109. https://doi.org/10.1177/1053815111418235

Machalicek, W., Gross, D. P., Armijo-Olivo, S., Ferriero, G., Kiekens, C., Martin, R., Walshe, M., & Negrini, S. (2024). The role of single case experimental designs in evidence creation in rehabilitation. European Journal of Physical and Rehabilitation Medicine, 60(6), 1100–1111. https://doi.org/10.23736/S1973-9087.24.08713-6

What Works Clearinghouse: the two-part bar

To meet WWC evidence standards, a study needs both an internally valid design and adequate interobserver agreement. Design rigor and measurement rigor are judged together; strength in one does not compensate for weakness in the other.

Note your weaknesses now, for your presentation

You are writing all of this into your proposal. Separately, when you present, name each place your design is weak and explain how and why it is a limitation. Identifying a limitation and accounting for it is a mark of rigor, not a flaw. Use the box below to start that list.

4Generalization & Maintenanceoptional

Generality refers to the extent to which a behavior change carries across settings, behaviors, materials, or people. You can choose to address generalization and maintenance by adding phases to your study, for example baseline, intervention, then a generalization phase, and a maintenance phase. What you plan here also feeds your external-validity argument in Section 3.

Sources you can cite for this section

Cite a source if you add generalization or maintenance. Options from your course readings:

Walker, S. G., & Carr, J. E. (2021). Generality of findings from single-case designs: It's not all about the N. Behavior Analysis in Practice, 14, 991–995. https://doi.org/10.1007/s40617-020-00547-3

Machalicek, W., Gross, D. P., Armijo-Olivo, S., Ferriero, G., Kiekens, C., Martin, R., Walshe, M., & Negrini, S. (2024). The role of single case experimental designs in evidence creation in rehabilitation. European Journal of Physical and Rehabilitation Medicine, 60(6), 1100–1111. https://doi.org/10.23736/S1973-9087.24.08713-6

Dallery, J., Cassidy, R. N., & Raiff, B. R. (2013). Single-case experimental designs to evaluate novel technology-based health interventions. Journal of Medical Internet Research, 15(2), e22. https://doi.org/10.2196/jmir.2227