What is AutoML
Automated Machine Learning (AutoML) aims to automate the parts of the machine learning pipeline that traditionally rely on human experience and repeated trial and error. It lets people without an algorithm background obtain usable models, and lets experts focus on problem definition and data quality.
What exactly gets automated
A full machine learning pipeline usually includes data cleaning and preprocessing, feature engineering, model selection, hyperparameter tuning, and finally ensembling and deployment. AutoML mainly covers model selection, hyperparameter optimization and preprocessing/feature pipeline search; some platforms extend to neural architecture search and model compression.
Common technical approaches
Hyperparameter optimization is dominated by Bayesian optimization (for example Gaussian-process or TPE based methods), which needs fewer evaluations than grid or random search. Other approaches include evolutionary algorithms, population-based training (PBT), and reinforcement learning for sequential decisions. Neural Architecture Search (NAS) treats the network structure itself as the search space — a related but heavier line of work.
Representative tools and platforms
In the open-source world, common options include Auto-sklearn, AutoKeras, TPOT, FLAML, and Optuna (focused on hyperparameter optimization). Cloud vendors also offer managed services, such as Google Cloud Vertex AI (formerly Cloud AutoML) and Azure Machine Learning Automated ML, which provide automated model building.
When it fits and when it does not
It fits structured tabular baseline modeling, relatively well-defined feature and hyperparameter spaces, and scenarios that value fast iteration. It fits less well when the problem itself is unclear, when the data has serious leakage or bias, or when there are extreme requirements on interpretability and inference latency — in those cases manual design is often more reliable.
Limitations and common misconceptions
AutoML does not raise the ceiling of model performance out of thin air; its value is approaching that ceiling faster while reducing human oversight. A poorly designed search space, evaluation metric or validation scheme can actually overfit the validation set. Compute cost, interpretability of results, and maintenance of automated pipelines in production are all practical concerns.
Sources: compiled from publicly available academic papers and industry materials, including documentation of open-source projects such as Auto-sklearn, TPOT and FLAML, and public AutoML product descriptions from cloud providers such as Google and Microsoft.