Skip to content
@trainindata

Train In Data

Advanced content creation for data science and machine learning

Welcome to Train in Data

GitHub followers LinkedIn

We are a group of passionate data scientists and software developers with the mission to make intermediate and advanced topics on machine learning, data science and AI software engineering accessible to the wider data science community.

We create intermediate and advanced online courses on machine learning, data science and AI software development. We also write books, and maintain open-source libraries for feature engineering and selection, including Feature-engine.

In addition, we talk, blog and participate in podcasts about machine learning, software development and open-source.

Online Courses

Check out the courses that we teach.

Courses What you will learn
Feature engineering for machine learning Learn to create new features, impute missing data, encode categorical variables, transform and discretize features and much more.
Feature selection for machine learning Learn to select features using wrapper, filter, embedded and hybrid methods, and build simpler and reliable models.
Master Hyperparameter Optimization for Tabular Learning Learn about grid and random search, Bayesian Optimization, Multi-fidelity models, Optuna, Hyperopt, Scikit-Optimize and more.
Machine learning with imbalanced data Learn about under- and over-sampling, ensemble and cost-sensitive methods and improve the performance of models trained on imbalanced data.
Feature engineering for time series forecasting Learn to create lag and window features, impute data in time series, encode categorical variables and much more, specifically for forecasting.
Forecasting with Machine Learning Learn to perform time series forecasting with machine learning models like linear regression, random forests and xgboost.
Machine Learning Interpretability Learn to interpret the predictions of your white box and black box machine learning models.
Clustering and Dimensionality Reduction Learn to extract information from unlabelled data through clustering and dimensionality reduction techniques.

Books

Find out more about machine learning through our books, and have the code at your fingertips.

Books Summary
Python feature engineering Cookbook, third edition Over 70 Python recipes to implement feature engineering in tabular, transactional, time series and text data.
Feature selection in machine learning, second edition Over 20 methods to select the most predictive features and build simpler, faster, and more reliable machine learning models.
Imbalanced Data: Myths, Mistakes and Modern Solutions A critical look at class imbalance, moving beyond SMOTE and default thresholds toward cost-sensitive learning, proper threshold tuning, and evaluation metrics that reflect real-world requirements.

Open-source

The open-source libraries we contribute to.

Library About Sponsor us
Feature-engine Multiple transformers for missing data imputation, categorical encoding, variable transformation and discretization, feature creation and more. Sponsor us
tsfresh Automatically create features for time series classification.
imbalanced-learn Tools for under- and over-sampling and dealing with imbalanced data.
BorutaPy Feature selection using Boruta.
Eli5 Tools for machine learning interpretability.

Our instructors

Get to know the creators and instructors of our courses.

Instructor Role
Soledad Galli Data scientist
Kishan Manani Data scientist
Dalibor Veljkovic Data scientist

Follow us

Follow us on social media or through our website to be up to date with our latest news.

Media Summary
Train in Data Enroll in our courses and books
Newsletter We talk about data science, machine learning and how to become a data scientist.
YouTube We post about data science, machine learning and how to become a data scientist.
LinkedIn We talk about data science, machine learning and how to become a data scientist.
BlueSky We post about data science, machine learning and how to become a data scientist.
Instagram We share bite-sized tips on data science and machine learning.
TikTok We share bite-sized tips on data science and machine learning.
Blog We write about data science, machine learning, feature engineering and selection and more.

Profile views counter


We hope to see you around.

Pinned Loading

  1. feature-engineering-for-time-series-forecasting feature-engineering-for-time-series-forecasting Public

    Code repository for the online course "Feature Engineering for Time Series Forecasting".

    Jupyter Notebook 206 136

  2. machine-learning-interpretability machine-learning-interpretability Public

    Forked from solegalli/machine-learning-interpretability

    Jupyter Notebook 7 7

  3. forecasting-with-machine-learning forecasting-with-machine-learning Public

    Code repository for the course "Forecasting with Machine Learning Models"

    Jupyter Notebook 31 16

Repositories

Showing 10 of 11 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…