The same engine,
a Pythonic API.
pip install millwright — a Pythonic pipeline over the same Rust engine, shipped on PyPI as an abi3 wheel built with maturin. Run it at Rust speed from a notebook.
A pipeline, from Python.
pip install millwright
import millwright as mw train = mw.Frame.from_pandas(df) # or from_numpy / from_rows pipe = (mw.Pipeline() .step("impute", mw.SimpleImputer.median()) .step("scale", mw.StandardScaler()) .estimator("rf", mw.RandomForest(n_trees=200, max_depth=8))) pipe.fit(train, y_train) preds = pipe.predict(test) metrics = pipe.evaluate(test, y_test) # -> {"accuracy": …, "f1": …}
The transformer / estimator objects (StandardScaler, MinMaxScaler, SimpleImputer, OneHotEncoder, RandomForest, LinearRegression, Knn, Svc, NaiveBayes) are the same engines as Rust. The older builder form — pipe.standard_scaler(), pipe.random_forest() — still works.
numpy, pandas, or a typed table.
# a Frame reads arrays and DataFrames directly train = mw.Frame.from_numpy(X) # or from_pandas(df) / from_rows(rows) # or the dtype-aware Table (strings, dates, nulls) + automated EDA data = mw.Table.from_csv("churn.csv") mw.Profile.of_with_target(data, "churned").to_html("eda.html") train = data.to_frame()
The whole lifecycle.
# grid search + stratified CV over the pipeline best = (mw.GridSearch(pipe, {"rf__max_depth": [4, 8, 16]}) .cv(mw.StratifiedKFold(5)).scoring("f1") .fit(train, y_train)) best.best_score; best.best_params() # manual ensembles and full AutoML are first-class APIs too another_pipe = mw.Pipeline().estimator("rf", mw.RandomForest(n_trees=120)) vote = (mw.Voting("hard", "classification") .add("rf1", pipe) .add("rf2", another_pipe)) vote.fit(train, y_train) soft_vote = (mw.Voting("soft", "classification") .add("lr1", mw.Pipeline().estimator("lr", mw.LogisticRegression())) .add("lr2", mw.Pipeline().estimator("lr", mw.LogisticRegression(l2=0.01)))) soft_vote.fit(train, y_train) probabilities = soft_vote.predict_proba(test) probability_pipe = mw.Pipeline().estimator("lr", mw.LogisticRegression()) probability_pipe.fit(train, y_train) probabilities = probability_pipe.predict_proba(test) auto = (mw.AutoML.classifier().budget_trials(40) .deployability("onnx") # use "any" for KNN/NB/SVC too .ensemble_kinds(["voting", "bagging", "boosting", "stacking"]) .fit(train, y_train)) rows = auto.leaderboard_entries() # [(config, score), …] failed = auto.candidate_failures() # candidates skipped safely refit_fallbacks = auto.refit_failures() # ranked winners that failed full refit winner = auto.best_model() # pipeline or ensemble auto.elapsed_seconds # measured search + refit time auto.completed_trials, auto.budget_exhausted # budget diagnostics if auto.supports_proba: probabilities = auto.predict_proba(test) # direct winner probabilities auto.export_onnx("automl.onnx") # SHAP importance, and one portable ONNX artifact pipe.fit(train, y_train) pipe.explain(test) # [(feature, mean|shap|), …] pipe.export_onnx("churn.onnx") # consume an external sklearn / PyTorch model (exported to ONNX) as a step ext = mw.Pipeline().estimator("onnx", mw.OnnxModel("model.onnx"))
python is deliberately not part of full: pyo3's extension-module defers libpython symbols, so a plain cargo test can't link it. It is built and tested the way it ships — as a wheel. To build from source, from a virtualenv: maturin develop --features python. The wheel bundles EDA, model selection, ensembles, AutoML, explainability, and ONNX.