
Tau U effect size
tau_u_computation.RdTau U effect size and variance computation.
Arguments
- data
Raw data to compute the effect sizes and variances from
- studyID
A character string representing the studyID
- subjectID
A character string representing the subjectID
- outcome_name
A character string representing the outcome name
- phase_name
A character string representing the phase name
- phase_order
Optional character vector of length 2 giving the baseline phase first and the comparison phase second. If
NULL, factor levels are used whenphase_nameis a factor; otherwise the two unique phase labels are sorted alphabetically.- version
Character string selecting the baseline-trend-corrected Tau-U denominator.
"revised"usesm * n;"original"usesm * n + m * (m - 1) / 2. Whenbaseline_trend_adjust = FALSE, the unadjusted A-vs-B Tau always usesm * n.- baseline_trend_adjust
Logical. If
TRUE, subtracts the baseline trend component from the phase contrast. IfFALSE, returns the unadjusted A-vs-B nonoverlap Tau.- variance_correction
Character string selecting an optional multiplicative correction for the three variance approximations.
"none"leavesv1,v2, andv3unchanged."small_sample"applies a finite-sample multiplier,"autocorrelation"applies an AR(1)-style lag-1 autocorrelation multiplier, and"both"applies both corrections.- na_option
Currently only listwise is supported
Value
Returns a data frame that contains the studyID, subjectID,
and the Tau-U effect size and 3 different variance approximations:
v1 is an empirical signed-comparison variance approximation under working
independence. v2 is the theoretical null variance of the signed phase
component, omitting baseline trend variability. v3 adds the baseline-trend
component when requested. The theoretical formulas assume independent,
continuous observations from a common distribution.
The output also includes the lag-1 autocorrelation estimate used by the
optional autocorrelation correction, the selected correction label, and the
final variance multiplier applied to v1, v2, and v3.
Details
Tau-U is computed by contrasting the first phase in phase_order with the
second. Supplying phase_order is recommended whenever the intended baseline
and intervention labels are known. Each study-subject combination must
contain observations from both phases. baseline_trend_adjust = TRUE
computes the baseline-trend-corrected Tau-U; baseline_trend_adjust = FALSE
computes the unadjusted A-vs-B nonoverlap Tau. For baseline-trend-corrected
Tau-U, the revised denominator uses m * n, while the original denominator
adds the baseline trend term m * (m - 1) / 2.
Let m be the number of baseline observations, n the number of comparison
observations, Q_P the vector of pairwise phase-contrast signs comparing
comparison observations to baseline observations, and Q_A the strictly
upper-triangular baseline self-comparison signs used for the baseline trend
adjustment. The Tau-U numerator is
$$
N_U = \sum Q_P - I_{\mathrm{trend}} \sum Q_A,
$$
where I_trend = 1 when baseline_trend_adjust = TRUE and 0 otherwise.
The reported Tau-U is
$$
\tau_U = N_U / d,
$$
where d = m n for the revised version and for all unadjusted A-vs-B Tau
values. The original baseline-trend-corrected version instead uses
d = m n + m (m - 1) / 2. The revised denominator keeps Tau-U on the
A-vs-B nonoverlap scale after removing baseline trend; the original
denominator normalizes by all pairwise comparisons entering the adjusted
numerator.
With fixed phase sizes and a fixed adjustment choice, v3 is the exact
variance of the stated statistic under the independent, continuous,
common-distribution null, before optional multipliers. Under that null,
v2 is exact for the phase component sum(Q_P) / d only; when adjustment is
enabled it omits uncertainty from baseline trend. Ties, phase differences,
temporal trends, and serial dependence depart from these assumptions; the
theoretical formulas do not include tie corrections. Data-dependent selection
of baseline adjustment requires a separate variance analysis.
The empirical variance approximation v1 is
$$
\widehat{\mathrm{Var}}(\tau_U) =
\frac{\widehat{\mathrm{Var}}(Q_P) \, (m n) +
\widehat{\mathrm{Var}}(Q_A) \, m (m - 1) / 2}{d^2},
$$
where d = m n for the revised version and
d = m n + m (m - 1) / 2 for the original baseline-trend-corrected version.
When baseline_trend_adjust = FALSE, the Q_A term is omitted and
d = m n. This is a working-independence approximation: it omits
covariance among comparisons sharing observations and between components.
Even independent observations produce dependent signs. For two phase signs
sharing a baseline observation, the covariance is 1/3 under the strict null.
Consequently, empirical sign dispersion times the comparison count need not
estimate the variance of the sum. v1 already uses signed comparisons and
requires no count-to-sign scaling correction.
The Mann-Whitney-style theoretical variance approximation v2 is
$$
\mathrm{Var}(\tau_U) =
\frac{m n (m + n + 1) / 3}{d^2},
$$
using the same denominator d.
The phase-contrast variance approximation v3 adds the baseline-trend
component when baseline_trend_adjust = TRUE:
$$
\mathrm{Var}_{\mathrm{pc}}(\tau_U) =
\frac{m n (m + n + 1) / 3 + m (m - 1) (2 m + 5) / 18}{d^2},
$$
again with d = m n for the revised version and
d = m n + m (m - 1) / 2 for the original baseline-trend-corrected version.
When baseline_trend_adjust = FALSE, the baseline-trend component is omitted,
so v2 and v3 are identical before optional corrections.
To derive the scale, write P = sum(Q_P) and Q = sum(Q_A).
Let U count cross-phase pairs with B_j > A_i, and W count baseline
pairs with A_k > A_i for i < k in chronological row order. Without ties,
$$P = 2U - mn, \qquad Q = 2W - \binom{m}{2}.$$
Under the strict null, each count indicator has variance 1/4. Cross-phase
indicators sharing one observation have covariance 1/12, giving
$$\mathrm{Var}(U) = mn/4 + 2[m\binom{n}{2} + n\binom{m}{2}]/12
= mn(m+n+1)/12.$$
For each baseline triple the three indicator covariances are
1/12, -1/12, 1/12, giving
$$\mathrm{Var}(W) = \binom{m}{2}/4 + 2\binom{m}{3}/12
= m(m-1)(2m+5)/72.$$
Disjoint comparisons are independent. Multiplication by two multiplies
variance by four, yielding signed-sum divisors 3 and 18 above.
In general the numerator variance is
$$\mathrm{Var}(P-IQ) = \mathrm{Var}(P) + I\mathrm{Var}(Q)
- 2I\mathrm{Cov}(P,Q),$$
where I = I_trend. Conditional on the unordered baseline values and all
comparison values, P is fixed and equally likely baseline time orders give
conditional mean zero for Q. Thus Cov(P,Q) = 0 under the strict null.
This argument does not establish exactness under the departures listed above.
Two optional corrections can be applied to all three approximate variances.
The small-sample correction multiplies the variances by
(m + n) / (m + n - 1). This is a heuristic for Tau-U; it inflates v3
above its exact strict-null variance. The autocorrelation
correction estimates the lag-1 autocorrelation rho from the observed
within-case outcome sequence, in the order rows appear in data, and applies
the AR(1)-style design-effect multiplier
$$
1 + 2 \sum_{k = 1}^{m+n-1} (1 - k / (m+n)) \rho^k.
$$
This finite-series factor describes a stationary AR(1) sample mean;
applying it to Tau-U is an approximation. Raw-outcome correlation can reflect
phase shifts or trends as well as serial dependence. It differs from the
infinite-series factor (1 + rho) / (1 - rho) in short series.
Both optional multipliers are applied after the signed-scale base variances
and do not make the formulas exact for other data-generating processes.
Polynomial calibration is a separate methodological task; coefficients fitted
to a different base-variance scale require separate reanalysis and validation.
The package bounds this design-effect multiplier above zero before applying
it. If autocorrelation cannot be estimated, the autocorrelation multiplier is
set to 1 and autocorrelation is returned as NA. With
variance_correction = "both", the final multiplier is the product of the
small-sample and autocorrelation multipliers:
$$
\widehat{\mathrm{Var}}_{\mathrm{corrected}}(\tau_U)
= c \widehat{\mathrm{Var}}(\tau_U).
$$
Examples
scd <- data.frame(
study = "S1",
subject = "P1",
phase = rep(c("A", "B"), each = 4),
outcome = c(2, 3, 3, 4, 5, 6, 6, 7)
)
tau_u_computation(
scd,
studyID = "study",
subjectID = "subject",
outcome_name = "outcome",
phase_name = "phase",
phase_order = c("A", "B")
)
#> study subject Tau_U v1 v2 v3 autocorrelation
#> S1||P1 S1 P1 0.6875 0.00390625 0.1875 0.2213542 0.9519231
#> variance_correction variance_multiplier
#> S1||P1 none 1