funcml: A Formula-First Framework for Machine Learning in R

One consistent API across 20+ learners, resampling, tuning, interpretation, and causal estimation

R
machine-learning
packages
Author

Imad El Badisy

Published

June 1, 2025

Machine learning in R has long suffered from a fragmentation problem. Fitting a random forest, a regularized linear model, and a neural network each requires learning a different package, a different formula syntax, and a different set of conventions for prediction and evaluation. funcml was built to close that gap.

What is funcml?

funcml provides a unified, formula-first interface to supervised learning in R. Six top-level verbs cover the full ML workflow:

Verb What it does
fit() Train any learner via a common formula interface
evaluate() Resampling-based performance with any metric
tune() Grid or random hyperparameter search
compare() Benchmark multiple models side by side
interpret() 14 model-agnostic interpretability methods
estimate() Plug-in g-computation for ATE/CATE estimation

Install from CRAN:

A minimal example

library(funcml)

# fit a random forest
fit_rf <- fit(Sepal.Length ~ ., data = iris, model = "ranger")

# evaluate with 5-fold CV
evaluate(data = iris, formula = Sepal.Length ~ ., model = "ranger", resampling = cv(5), seed = 1)
<funcml_eval> model: ranger | task: regression
  metric   mean     sd n std_error conf_level conf_low conf_high
1   rmse 0.3391 0.0451 5    0.0202       0.95   0.2830    0.3951
2    mae 0.2771 0.0303 5    0.0136       0.95   0.2394    0.3147
3    mse 0.1166 0.0322 5    0.0144       0.95   0.0767    0.1565
4  medae 0.2428 0.0391 5    0.0175       0.95   0.1943    0.2913
5   mape 0.0478 0.0050 5    0.0022       0.95   0.0416    0.0540
6    rsq 0.8236 0.0552 5    0.0247       0.95   0.7550    0.8922
# compare to a regularized regression
compare(
  data = iris,
  formula = Sepal.Length ~ .,
  models = c("ranger", "glmnet"),
  resampling = cv(5),
  seed = 1
)
<funcml_compare> task: regression | tuned: FALSE
    model metric    mean     sd n std_error conf_level conf_low conf_high tuned
1  ranger   rmse  0.3391 0.0451 5    0.0202       0.95   0.2830    0.3951 FALSE
2  ranger    mae  0.2771 0.0303 5    0.0136       0.95   0.2394    0.3147 FALSE
3  ranger    mse  0.1166 0.0322 5    0.0144       0.95   0.0767    0.1565 FALSE
4  ranger  medae  0.2428 0.0391 5    0.0175       0.95   0.1943    0.2913 FALSE
5  ranger   mape  0.0478 0.0050 5    0.0022       0.95   0.0416    0.0540 FALSE
6  ranger    rsq  0.8236 0.0552 5    0.0247       0.95   0.7550    0.8922 FALSE
7  glmnet   rmse  0.8258 0.0670 5    0.0300       0.95   0.7426    0.9090 FALSE
8  glmnet    mae  0.6914 0.0683 5    0.0305       0.95   0.6066    0.7762 FALSE
9  glmnet    mse  0.6856 0.1110 5    0.0496       0.95   0.5478    0.8234 FALSE
10 glmnet  medae  0.6252 0.0687 5    0.0307       0.95   0.5399    0.7104 FALSE
11 glmnet   mape  0.1205 0.0104 5    0.0047       0.95   0.1076    0.1335 FALSE
12 glmnet    rsq -0.0175 0.0208 5    0.0093       0.95  -0.0433    0.0084 FALSE
   rank
1     1
2     1
3     1
4     1
5     1
6     1
7     2
8     2
9     2
10    2
11    2
12    2

Interpretability built in

interpret() wraps 14 methods, from permutation importance and partial dependence plots to SHAP values and accumulated local effects, behind a single call:

interpret(fit_rf, data = iris, method = "pdp", features = "Petal.Width")$result$curves |> head()
      feature     value     yhat
1 Petal.Width 0.1000000 5.567361
2 Petal.Width 0.2142857 5.585780
3 Petal.Width 0.3285714 5.595121
4 Petal.Width 0.4428571 5.676712
5 Petal.Width 0.5571429 5.669333
6 Petal.Width 0.6714286 5.680577

Causal estimation

When the research question goes beyond prediction, estimate() wraps plug-in g-computation to estimate average and conditional treatment effects:

set.seed(1)
n <- 200
X1 <- rnorm(n); X2 <- rnorm(n)
A <- rbinom(n, 1, plogis(0.3 * X1))
Y <- 1 + 0.5 * A + 0.4 * X1 - 0.2 * X2 + rnorm(n)
mydata <- data.frame(Y, A, X1, X2)

# ATE of treatment A on outcome Y, adjusted for covariates X
estimate(
  data      = mydata,
  formula   = Y ~ A + X1 + X2,
  model     = "ranger",
  treatment = "A",
  estimand  = "ATE",
  seed      = 1
)
<funcml_estimand> ATE via g-computation
Treatment: A (1 vs 0)
Estimate: 0.4531 | SE: 0.0156 | 95% normal CI [0.4225, 0.4836]

The learner registry

funcml ships with 26 learners spanning linear models, tree ensembles, SVMs, neural networks, additive models, boosting, and stacking. All are accessed by name:

list_learners()[, c("learner", "has_tune")]
        learner has_tune
17     adaboost     TRUE
24         bart     TRUE
11          C50     TRUE
20      cforest     TRUE
19        ctree     TRUE
7      densemlp     TRUE
8     e1071_svm     TRUE
13        earth     TRUE
16          fda     TRUE
14          gam     TRUE
10          gbm     TRUE
1           glm     TRUE
3        glmnet     TRUE
12         kknn     TRUE
21          lda     TRUE
23     lightgbm     TRUE
6           mlp     TRUE
15   naivebayes     TRUE
5          nnet     TRUE
18          pls     TRUE
22          qda     TRUE
9  randomForest     TRUE
4        ranger     TRUE
2         rpart     TRUE
26     stacking     TRUE
27 superlearner     TRUE
25      xgboost     TRUE

Citation

@software{El_Badisy_funcml_2026,
  author  = {El Badisy, Imad},
  license = {GPL-3.0-only},
  month   = apr,
  title   = {{funcml}},
  doi     = {10.5281/zenodo.20707605},
  url     = {https://doi.org/10.5281/zenodo.20707605},
  version = {0.7.1},
  year    = {2026}
}
Back to top