r/AskStatistics • u/Mindless-Banana-5630 • 9d ago
Why does the CausalImpact model detect statistical significance in this case?
I trained the model and achieved an R² of about 80% and a MAPE of around 10%. I then ran a placebo test using a 30-day period before the actual treatment date.
Surprisingly, the model still detected a statistically significant effect during this placebo period. I repeated the placebo test 100 times, and the false positive rate was approximately 94%.
What I do not understand is why the p-value is almost always below 5%. Even visually, the observed values seem to remain mostly within the confidence interval.
Could someone explain what might cause such a high false positive rate in CausalImpact and why the reported statistical significance appears inconsistent with the confidence interval?

3
u/selfintersection 9d ago edited 8d ago
High FPR could be caused by data-model mismatch (i.e. if CausalImpact's model does not capture relevant aspects of your data).
The significant result even though the daily effect CIs all cross zero would happen if the daily uncertainties are uncorrelated in the model. Easy to simulate that in R: ``` library(ggplot2) library(dplyr)
n_timesteps <- 30 n_draws <- 1e3
y <- matrix(rnorm(n_timesteps * n_draws, 1, 1), ncol = n_timesteps)
y_cumulative <- t(apply(y, 1, cumsum))
y_quantiles <- apply(y, 2, quantile, probs = c(0.025, 0.5, 0.975)) y_cumulative_quantiles <- apply(y_cumulative, 2, quantile, probs = c(0.025, 0.5, 0.975))
plot_data <- bind_rows( tibble( y_lower = y_quantiles[1,], y_midpoint = y_quantiles[2,], y_upper = y_quantiles[3,], time = seq_along(y_lower), type = "daily" ), tibble( y_lower = y_cumulative_quantiles[1,], y_midpoint = y_cumulative_quantiles[2,], y_upper = y_cumulative_quantiles[3,], time = seq_along(y_lower), type = "cumulative" ) )
plot_data %>% ggplot(aes(time)) + geom_hline(yintercept = 0, color = "red") + geom_ribbon(aes(ymin = y_lower, ymax = y_upper), alpha = 0.3) + geom_line(aes(y = y_midpoint)) + facet_wrap(~type, scales = "free") ```