mimar: Compact Multiple Imputation, Assessment, and Reporting in R

Artificial amputation, chained-equations MI with ML imputers, diagnostic evaluation, and pooling, in one package

R
missing-data
packages
Author

Imad El Badisy

Published

August 15, 2025

Missing data is the rule, not the exception, in health research. Yet most workflows still string together three or four packages: one to characterize missingness, one to impute, one to pool, one to evaluate, with no consistent API across them. mimar brings all of that under one roof.

What is mimar?

mimar (v0.8.0) covers the full missing-data analysis cycle:

  1. Amputation: artificially introduce missingness into a complete dataset for simulation studies
  2. Imputation: single or multiple imputation via chained equations with a choice of statistical or ML imputers
  3. Evaluation: diagnostic metrics comparing imputed distributions to observed and (when known) true values
  4. Pooling: Rubin’s rules applied to post-imputation model estimates

Install:

Artificial amputation

ampute() introduces controlled missingness into a complete data frame, specifying the mechanism (MCAR, MAR, MNAR) and the proportion of missing cells:

library(mimar)

set.seed(1)
amp <- ampute(iris, prop = 0.25, mechanism = "MCAR")

head(amp$data)       # the amputed data frame
   Sepal.Length Sepal.Width Petal.Length Petal.Width Species
          <num>       <num>        <num>       <num>  <fctr>
1:          5.1         3.5          1.4         0.2  setosa
2:          4.9         3.0           NA         0.2  setosa
3:          4.7         3.2          1.3         0.2    <NA>
4:          4.6         3.1          1.5          NA  setosa
5:           NA         3.6          1.4          NA  setosa
6:          5.4          NA          1.7         0.4  setosa

Multiple imputation

impute() runs chained equations with any registered imputer. The default is predictive mean matching ("pmm"), but ML-based imputers are equally accessible:

# PMM: the safe default
imp_pmm <- impute(amp, m = 5, imputer = "pmm", maxit = 5)

# Random forest imputation
imp_rf <- impute(amp, m = 5, imputer = "ranger")

# XGBoost imputation
imp_xgb <- impute(amp, m = 5, imputer = "xgboost")

class(imp_pmm)
[1] "mimar_imputation" "list"            

impute() returns a mimar_imputation object containing all m completed datasets alongside metadata for downstream steps.

Available imputers

imputer_registry()[, c("imputer", "implementation", "stochastic")]
         imputer implementation stochastic
          <char>         <char>     <lgcl>
 1:         mean          mimar      FALSE
 2:       median          mimar      FALSE
 3:         mode          mimar      FALSE
 4:        naive          mimar      FALSE
 5:         norm          mimar       TRUE
 6:          pmm          mimar       TRUE
 7:         spmm          mimar       TRUE
 8:       logreg          mimar       TRUE
 9:      polyreg          mimar       TRUE
10:           rf        wrapped       TRUE
11:       ranger        wrapped       TRUE
12:        rpart        wrapped       TRUE
13:       nbayes        wrapped       TRUE
14:          svm        wrapped       TRUE
15:         bart        wrapped       TRUE
16:       glmnet        wrapped       TRUE
17:          gbm        wrapped       TRUE
18:      xgboost        wrapped       TRUE
19:          knn          mimar       TRUE
20:      hotdeck          mimar       TRUE
21:         famd        wrapped       TRUE
22: superlearner          mimar       TRUE
23:           sl          mimar       TRUE
24:     densemlp        wrapped       TRUE
25:      missknn        wrapped       TRUE
         imputer implementation stochastic
          <char>         <char>     <lgcl>

Evaluation

When the true values are known (after amputation, or from a simulation), evaluate() computes recovery metrics comparing imputed to true values:

diag <- evaluate(imp_pmm)
diag$distribution       # observed vs. imputed means, SDs, unique counts
       variable    type missing observed_unique imputed_unique observed_mean
         <char>  <char>   <int>           <int>          <int>         <num>
1: Sepal.Length numeric      33              33             28      5.837607
2:  Sepal.Width numeric      39              22             19      3.075676
3: Petal.Length numeric      34              42             38      3.755172
4:  Petal.Width numeric      34              21             21      1.252586
5:      Species  factor      31               3              3            NA
   imputed_mean observed_sd imputed_sd
          <num>       <num>      <num>
1:     5.987273   0.8377645  0.7791618
2:     3.035385   0.4509014  0.4011957
3:     3.770000   1.8037808  1.5557433
4:     1.076471   0.7519784  0.7701759
5:           NA          NA         NA

Completing and pooling

Extract a single completed dataset or pool regression estimates across all m imputations using Rubin’s rules:

completed <- complete(imp_pmm, which = 1)
head(completed)
   Sepal.Length Sepal.Width Petal.Length Petal.Width    Species
          <num>       <num>        <num>       <num>     <fctr>
1:          5.1         3.5          1.4         0.2     setosa
2:          4.9         3.0          1.4         0.2     setosa
3:          4.7         3.2          1.3         0.2 versicolor
4:          4.6         3.1          1.5         0.2     setosa
5:          5.1         3.6          1.4         0.5     setosa
6:          5.4         3.3          1.7         0.4     setosa
# fit a model on each completed dataset, then pool the coefficients
completed_list <- lapply(seq_len(imp_pmm$m), function(k) complete(imp_pmm, which = k))
fits  <- lapply(completed_list, function(d) lm(Sepal.Length ~ Sepal.Width + Petal.Length + Petal.Width, data = d))
betas <- lapply(fits, coef)
vars  <- lapply(fits, function(f) diag(vcov(f)))

pool(betas, variance = vars)
mimar pooled results
           term   estimate  std.error statistic    df      p.value   conf.low
         <char>      <num>      <num>     <num> <num>        <num>      <num>
1:  (Intercept)  2.1204880 0.23942658  8.856527   Inf 8.254583e-19  1.6512205
2:  Sepal.Width  0.5861945 0.06375298  9.194779   Inf 3.757668e-20  0.4612410
3: Petal.Length  0.7034397 0.05719104 12.299822   Inf 9.077385e-35  0.5913473
4:  Petal.Width -0.5639728 0.12879778 -4.378746   Inf 1.193641e-05 -0.8164118
    conf.high     m within_variance between_variance total_variance
        <num> <int>           <num>            <num>          <num>
1:  2.5897555     5     0.057325087                0    0.057325087
2:  0.7111480     5     0.004064442                0    0.004064442
3:  0.8155320     5     0.003270815                0    0.003270815
4: -0.3115338     5     0.016588867                0    0.016588867
   relative_increase_variance   rule
                        <num> <char>
1:                          0  rubin
2:                          0  rubin
3:                          0  rubin
4:                          0  rubin

When to use mimar

mimar is designed for:

  • Simulation studies: amputate, impute, evaluate recovery under controlled conditions
  • Applied health research: handle missing covariates or outcomes before regression or survival analysis
  • ML-driven imputation: when the imputation model is complex enough to warrant a tree ensemble or gradient booster
Back to top