Skip to main content

Machine learning

scikit-learn datasets, models, evaluation and tuning.

15 nodes. Right-click any node in the editor to read this documentation in the app.

ML / Dataโ€‹

๐Ÿ”„ Apply Transformerโ€‹

id sklearn_transform ยท ML / Data ยท Python export: yes

Fit a transformer on the TRAINING rows only, then transform both training and test rows (so nothing about the test data leaks into preprocessing). Outputs a new splits dict; y is unchanged.

Inputs

PortTypeDescription
transformerTransformerUnfitted transformer from a Transformer node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
splitsSplitsSplits with X_train / X_test transformed.

๐ŸŒธ Sample Datasetโ€‹

id sklearn_load_dataset ยท ML / Data ยท Python export: yes

One of scikit-learn's built-in datasets as a DataFrame: the feature columns plus a 'target' column. iris, wine, breast_cancer and digits are classification problems; diabetes is regression. Ships with scikit-learn, so it works offline.

Outputs

PortTypeDescription
dataframeDataFrameFeature columns plus a 'target' column.

Fields

FieldTypeDefaultChoices
Datasetselectirisiris, wine, breast_cancer, diabetes, digits

โœ‚๏ธ Train/Test Splitโ€‹

id sklearn_train_test_split ยท ML / Data ยท Python export: yes

Split a DataFrame into training and test sets. Name the target column; every other column becomes a feature. A fixed random seed makes the split reproducible (use a negative number for a different split each run); 'Stratify' keeps class proportions equal in both sets.

Inputs

PortTypeDescription
dataframeDataFrameTable with the features and the target column.

Outputs

PortTypeDescription
splitsSplitsDict of X_train, X_test, y_train, y_test (pandas objects).

Fields

FieldTypeDefaultChoices
Target columntexttarget
Test fractionfloat0.2
Random seed (negative = random)int42
Stratify by targetcheckboxfalse
Shuffle before splittingcheckboxtrue

๐Ÿ”ง Transformerโ€‹

id sklearn_transformer ยท ML / Data ยท Python export: yes

Configure a preprocessing step: standard / min-max / robust scaling, missing-value imputation (mean, median, most frequent), or PCA. Nothing is fitted here โ€” wire it into Apply Transformer. 'Components' applies to PCA only; 'Extra parameters' is an optional JSON object of further scikit-learn arguments.

Outputs

PortTypeDescription
transformerTransformerUnfitted scikit-learn transformer.

Fields

FieldTypeDefaultChoices
Transformerselectstandard_scalerstandard_scaler, minmax_scaler, robust_scaler, impute_mean, impute_median, impute_most_frequent, pca
Components (PCA only)int2
Extra parameters (JSON)text

ML / Evaluationโ€‹

๐Ÿ”ฒ Confusion Matrixโ€‹

id sklearn_confusion_matrix ยท ML / Evaluation ยท Python export: yes

For a fitted classifier: how often each actual class was predicted as each class (rows = actual, columns = predicted). The diagonal is the correct predictions.

Inputs

PortTypeDescription
modelFittedModelA fitted model โ€” from Fit or a Hyper-parameter Search node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
matrixDataFrameCounts, rows = actual class, columns = predicted class.

Fields

FieldTypeDefaultChoices
Rowsselecttesttest, train

๐Ÿ” Cross-Validateโ€‹

id sklearn_cross_val ยท ML / Evaluation ยท Python export: yes

k-fold cross-validation of an (unfitted) estimator on the TRAINING rows: a more reliable estimate than a single split. Outputs one score per fold; scoring 'default' uses the model's own (accuracy for classifiers, Rยฒ for regressors).

Inputs

PortTypeDescription
estimatorEstimatorUnfitted model from a Classifier, Regressor or Clusterer node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
scoresDataFrameOne row per fold: fold, score.

Fields

FieldTypeDefaultChoices
Foldsint5
Scoringselectdefaultdefault, accuracy, f1_weighted, roc_auc, r2, neg_mean_squared_error, neg_mean_absolute_error

๐Ÿ“Š Evaluateโ€‹

id sklearn_evaluate ยท ML / Evaluation ยท Python export: yes

Score a fitted model. Classifiers: accuracy, precision, recall and F1 (class-weighted) plus ROC-AUC when the model gives probabilities. Regressors: Rยฒ, MAE, MSE and RMSE. Outputs a metric / value table (plot it with Bar Chart).

Inputs

PortTypeDescription
modelFittedModelA fitted model โ€” from Fit or a Hyper-parameter Search node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
metricsDataFrameTwo columns: metric, value.

Fields

FieldTypeDefaultChoices
Rowsselecttesttest, train

โญ Feature Importanceโ€‹

id sklearn_feature_importance ยท ML / Evaluation ยท Python export: yes

Which features drive a fitted model? Uses the model's feature_importances_ (forests, boosting, trees) or the magnitude of its coefficients (linear models, averaged over classes). Sorted, most important first.

Inputs

PortTypeDescription
modelFittedModelA fitted model โ€” from Fit or a Hyper-parameter Search node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
importanceDataFrameColumns: feature, importance.

ML / Modelsโ€‹

๐Ÿง  Classifierโ€‹

id sklearn_classifier ยท ML / Models ยท Python export: yes

Configure a classification model: logistic regression, random forest, gradient boosting, SVM, k-NN or decision tree. Nothing is trained until Fit. Only the parameters that apply to the chosen model are used: 'Trees' โ†’ random forest / gradient boosting; 'Max depth' โ†’ forests, boosting, decision tree (0 = no limit); 'C' โ†’ logistic regression / SVM; 'Neighbours' โ†’ k-NN; 'Learning rate' โ†’ gradient boosting; 'Kernel' โ†’ SVM. 'Extra parameters' is an optional JSON object of further scikit-learn arguments, e.g. {"min_samples_leaf": 3}.

Outputs

PortTypeDescription
estimatorEstimatorUnfitted scikit-learn model โ€” wire into Fit, Cross-Validate or Search.

Fields

FieldTypeDefaultChoices
Modelselectrandom_forestlogistic_regression, random_forest, gradient_boosting, svm, knn, decision_tree
Treesint100
Max depth (0 = no limit)int0
C (regularisation)float1.0
Neighboursint5
Learning ratefloat0.1
Kernelselectrbfrbf, linear, poly, sigmoid
Random seed (negative = random)int42
Extra parameters (JSON)text

๐Ÿซง Clustererโ€‹

id sklearn_clusterer ยท ML / Models ยท Python export: yes

Configure an unsupervised clustering model: k-means, DBSCAN, agglomerative or Gaussian mixture. 'Clusters' applies to k-means, agglomerative and the mixture; 'Eps' / 'Min samples' to DBSCAN. Fit it on the features (labels are ignored), then use Predict.

Outputs

PortTypeDescription
estimatorEstimatorUnfitted scikit-learn model โ€” wire into Fit, Cross-Validate or Search.

Fields

FieldTypeDefaultChoices
Algorithmselectkmeanskmeans, dbscan, agglomerative, gaussian_mixture
Clustersint3
Eps (DBSCAN)float0.5
Min samples (DBSCAN)int5
Random seed (negative = random)int42
Extra parameters (JSON)text

๐Ÿ‹๏ธ Fitโ€‹

id sklearn_fit ยท ML / Models ยท Python export: yes

Train a model on the training rows (X_train, and y_train when the model is supervised). Works on a copy, so the estimator node can feed other nodes unchanged. Outputs the fitted model.

Inputs

PortTypeDescription
estimatorEstimatorUnfitted model from a Classifier, Regressor or Clusterer node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
modelFittedModelThe trained model โ€” wire into Predict, Evaluate, Feature Importance.

๐Ÿ”ฎ Predictโ€‹

id sklearn_predict ยท ML / Models ยท Python export: yes

Run a fitted model on the test (or training) rows and output its predictions as a Series indexed like the rows. For clusterers these are the cluster labels; DBSCAN and agglomerative clustering can only label their training rows.

Inputs

PortTypeDescription
modelFittedModelA fitted model โ€” from Fit or a Hyper-parameter Search node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
predictionsSeriesOne prediction per row, named 'prediction'.

Fields

FieldTypeDefaultChoices
Rowsselecttesttest, train

๐Ÿ“ˆ Regressorโ€‹

id sklearn_regressor ยท ML / Models ยท Python export: yes

Configure a regression model: linear, ridge, lasso, random forest, gradient boosting, SVR, k-NN or decision tree. Nothing is trained until Fit. 'Alpha' applies to ridge / lasso. Only the parameters that apply to the chosen model are used: 'Trees' โ†’ random forest / gradient boosting; 'Max depth' โ†’ forests, boosting, decision tree (0 = no limit); 'C' โ†’ logistic regression / SVM; 'Neighbours' โ†’ k-NN; 'Learning rate' โ†’ gradient boosting; 'Kernel' โ†’ SVM. 'Extra parameters' is an optional JSON object of further scikit-learn arguments, e.g. {"min_samples_leaf": 3}.

Outputs

PortTypeDescription
estimatorEstimatorUnfitted scikit-learn model โ€” wire into Fit, Cross-Validate or Search.

Fields

FieldTypeDefaultChoices
Modelselectrandom_forestlinear_regression, ridge, lasso, random_forest, gradient_boosting, svr, knn, decision_tree
Alpha (ridge / lasso)float1.0
Treesint100
Max depth (0 = no limit)int0
Neighboursint5
Learning ratefloat0.1
C (SVR)float1.0
Kernel (SVR)selectrbfrbf, linear, poly, sigmoid
Random seed (negative = random)int42
Extra parameters (JSON)text

ML / Tuningโ€‹

id sklearn_search ยท ML / Tuning ยท Python export: yes

Find the best settings by cross-validation on the TRAINING rows. Give a JSON grid of parameter names to lists of values, e.g. {"n_estimators": [50, 100], "max_depth": [3, 5, 10]} โ€” use the scikit-learn parameter names of the estimator. 'grid' tries every combination; 'random' tries a random sample of them. The output is the fitted search: it acts as the best model (wire it into Predict / Evaluate) and into Search Results for the full table.

Inputs

PortTypeDescription
estimatorEstimatorUnfitted model from a Classifier, Regressor or Clusterer node.
splitsSplitsTrain/test data โ€” from Train/Test Split or Apply Transformer.

Outputs

PortTypeDescription
searchFittedModelFitted search; predicts with the best parameters found.

Fields

FieldTypeDefaultChoices
Strategyselectgridgrid, random
Parameter grid (JSON)textarea{"n_estimators": [50, 100], "max_depth": [3, 5,โ€ฆ
CV foldsint5
Scoringselectdefaultdefault, accuracy, f1_weighted, roc_auc, r2, neg_mean_squared_error, neg_mean_absolute_error
Samples (random strategy)int10
Random seed (negative = random)int42

๐Ÿ† Search Resultsโ€‹

id sklearn_search_results ยท ML / Tuning ยท Python export: yes

Every parameter combination a search tried, with its mean and spread of cross-validation score and rank โ€” best first.

Inputs

PortTypeDescription
searchFittedModelOutput of a Hyper-parameter Search node.

Outputs

PortTypeDescription
resultsDataFrameOne row per combination: parameters, mean_test_score, std_test_score, rank_test_score.