A machine learning roadmap for beginners in 2026

I wrote my data science roadmap 8 years ago. Some of its advice is still actual: learn programming, understand statistics, complete projects, and practice working with data. But I would change the emphasis now, especially with AI tools able to write much of the code.
For example, in 2018 I wrote that you should develop a skill of formulating thoughts, searching for information, finding it and understanding it. Finding answers is much easier now, since you can just ask AI. But with the abundance of information and the ease of getting it, checking whether the answer is correct has become even more important, and it still requires learning the fundamentals and getting experience.
In August, I gave a talk about building an ML career and answered questions from students. Many were worried about the same things: what to study, how much mathematics they need, which projects are worth doing, and whether their skills will become outdated before they get a job. I often get asked similar questions in DMs, so I wanted to write a blogpost with an updated advice.
My perspective comes mainly from applied ML work, so most of my suggestions will focus on learning concepts and applying them in practice.
Build foundations you can use

First, decide what kind of work you want to do. A data scientist focuses on experiments, an MLE - on training and deploying models, and a researcher - on developing new architectures. Nowadays, with the help of AI, some of the boundaries can be blurred, but the work still differs enough.
For an applied ML role, I suggest these foundations:
- Python: functions, data structures, reading files, working with libraries, and debugging errors. Get comfortable reading someone else’s code and making changes to it.
- Working with data: SQL queries, pandas or a similar library, basic plots for Exploratory Data Analysis. Lead how to deal with messy data.
- Mathematics and statistics: vectors and matrices, derivatives, probability, sampling, and uncertainty. Learn just enough to understand how ML algorithms and hypothesis testing work. You can learn more if you want to.
- ML fundamentals: regression and classification, simple baselines, overfitting, validation approaches, and choosing metrics. Understand how to evaluate model performance honestly.
- Basic software development: Git, environments and dependencies, making sure other people can run your code.
You can learn these together: train a small model, study the underlying algorithm, evaluate the results and publish to GitHub. Research requires depth in mathematics, but applied work needs just enough to judge whether a result is correct.
If you prefer a structured approach, I suggest going through 1-2 courses. There is no need to study more - don’t fall into a trap of endlessly taking courses and never doing a project by yourself. I don’t keep up with the modern courses, but I can recommend Andrew Ng’s Machine Learning Specialization. I also still like mlcourse.ai for practice.
I suggest studying Deep learning and LLMs later, after you get familiar with basics. And don’t study them in a vacuum - combine the learning with working on projects. For example, building a document search application gives you a reason to learn about embeddings and retrieval, and you don’t need to study every model architecture before starting it.
Let each project have a purpose

I distinguish between two kinds of projects: projects for learning and projects for showing to other people. They can overlap, but they don’t have to.
A learning project can be based on a tutorial, it can be straightforward, and it can be small. Its purpose is to teach you something new, and it doesn’t have to be impressive. The important part is implementing something by yourself, doing every step with a purpose, dealing with errors, and understanding the results. It is worth doing even if you never put the project on your CV.
A portfolio project needs to make your work understandable and presentable to someone else. Copying a tutorial isn’t enough - it doesn’t demonstrate your skills or your ability to make decisions. When I was starting my ML career, many beginner portfolios had a model for the Titanic or Iris dataset; later I saw a lot of Kaggle notebooks with a lot of visualizations and no analysis at all; now there are dozens of nearly identical to-do apps built with coding agents. Choose a problem you care about, for example, working with data from a hobby. A familiar problem with your own constraint is enough; you don’t have to search for a truly novel idea.
One of my early projects was about recognizing handwritten digits. Training the model was only part of it: I wanted people to be able to draw a digit on their phone and see a prediction. I knew very little about web development, spent days figuring out the drawing canvas, and learned how to pass its input to the model. During interviews, I could hand someone my phone and let them play with it.
That demo isn’t impressive today, but it was enough to show that I could complete a project and learn new skills. Now you can find a more novel idea to demonstrate your skills, but the principle is the same.
For a project you plan to share, include a clear README, a way to run it, your evaluation, and examples of failures. Interactive projects are better than static ones because people can try them, but a good analysis works too.
Understand the problem before building a model

Before spending weeks improving a model, find out how its predictions would be used.
Many years ago, during my consulting work, we considered forecasting a department’s budget. The historical data seemed suitable, so we started building a time-series model. However, the discussions with the department revealed that the planned budget was simply increased by 10% each year and the department always made sure that actual spendings matched the plan by the end of the year. It turned out there was no forecasting problem to solve.
On a beginner project, a similar case would be talking to the person who would use your application and checking whether a simple script or a few rules could be enough to solve their problem. If yes, use the simpler solution: knowing when not to use machine learning is part of the job. You can still train a model separately if your purpose is to learn.
Once you have a model, evaluation deserves at least as much attention as training; I consider it the most important skill for an ML engineer. Separate the data used for learning from the data used for checking results, and think about how the system will encounter new examples. Predicting next month’s demand requires a different split from randomly dividing an unrelated collection of images. The scikit-learn guide to common pitfalls is a good introduction to data leakage (information entering training that would not be available when making real predictions).
Look at individual mistakes as well as the average score. At Careem, we found large groups of supposedly matching driver photos in a facial recognition project. It turned out that some photos had been taken in offices with the same portrait on the wall, and the system was detecting that face. Inspecting the images explained a result that looked very strange in the output.
The consequences of errors mattered too: incorrectly blocking a legitimate driver was unacceptable. In another project, a prediction had to be done in under a second. For your own application, check what a wrong answer costs the user, how long they wait, and whether the system can decline to answer when it lacks enough information.
Use AI tools and learn enough to check them

I use coding agents for my own projects and competitions, and they save a lot of implementation time: in the ROGII competition, they wrote almost all of my code. But I still need to decide which experiments make sense and inspect the results. I have had to stop agents that kept repeating an unproductive approach or made incorrect assumptions or invented erroneous reasons of their failures.
In one experiment, an agent reported a large improvement in model quality. It had included the target variable among the input features. The code ran and the metric improved, but the experiment was invalid. Recognizing that mistake required understanding the data and the evaluation setup.
For a beginner, I would suggest using AI as a learning tool. When studying, ask AI to coach you, to ask you questions and to avoid giving you direct answers. When writing code, ask it to review your implementation, to explain errors, and to suggest improvements. But don’t let it do the thinking for you.
One great benefit of AI is that it can suit your learning style. Whether you prefer to read, watch, or listen, you can find explanations in your preferred format. You can also ask it to explain a concept in different ways until you understand it.
Use Kaggle and open source for additional practice

Kaggle helped me develop my skills and meet people, and I still participate in competitions (although rarely). It is especially good for learning to run experiments quickly, validate models, and investigate why an idea failed. Reading the writeups of the winners after a competition shows you approaches you would not have considered yourself.
But I wouldn’t recommend going pushing for a high leaderboard ranking as a reliable career plan. You would spend hundreds of hours without any guarantee of a medal, and the time spent on a competition is time not spent on other projects or learning. More than that, nowadays only a handful of companies value Kaggle medals in hiring, and they are usually looking for a very specific skill set. And winning in competitions often requires applying tricks and tweaks that are not relevant to real-world problems.
Participate in competitions with a learning goal, for example, to develop a skill of working with time-series. Build your own baseline and compare it with public solutions. A medal is a nice result, but you don’t need one to get an ML job.
Open source gives you practice working within someone else’s project - just like at work projects. Libraries you already use and like are a good place to begin, and there are several ways to contribute:
- Fix a problem you encountered while using a library or provide a small example that reproduces it.
- Look for a
good first issuethat fits your current skills. GitHub’s contribution guide explains how to find these and check repository activity. - Explore a smaller, less developed repository where a missing example, test, or documentation section would help its users.
Start small and read the contribution instructions. A smaller repository is often easier to understand, but check whether its maintainers are active before investing much time. Reproduce the issue, discuss substantial changes, and submit a fix that you can explain. Don’t just use AI to create a pull request end-to-end. First of all, you’ll learn nothing. And the maintainers are often flooded with low-quality PRs, so you may not get a response.
Going through a review and responding to feedback is a large part of what you learn. My first pull request to albumentations was closed because the library already had a similar transform. The second one, which added the GlassBlur augmentation, was merged after reviewers suggested a faster implementation and asked me to fix a failing test.
Read papers to understand the ideas behind the code

Reading papers is a good way to understand the ideas behind the code you are using. It is also a good way to learn how to evaluate a method and how to compare it with others.
If you are interested in NLP, Word2Vec, the Transformer, and BERT are good papers to get started with. They give you concrete examples of learning word representations, using attention, and pretraining with context from both sides of a word. Read them together with explanations and implementations, and don’t expect to understand every equation on the first attempt. There area many websites that share paper implementations, I can recommend labml.
For each paper, try to understand the problem, the proposed change, and the evidence that it is useful. I have written more than 200 paper reviews, and I follow a similar structure in them. Pay attention to the comparison: a method improving over one baseline on one dataset does not establish that it will generalize well. Reading the limitations and inspecting the code are often as useful as reading the announced result.
You can try to reproduce the results of a paper, but in many cases it isn’t possible. Some papers require a lot of data or computing resources, and some papers don’t provide enough details to reproduce the results.
For applied engineering, I usually start with papers related to a problem in my current project. For example, when I was working on a medical chat-bot project, I read papers about medical question answering and retrieval.
Expect the tools and approaches to change
Some of what you learn now will become outdated in several yers, and that is normal. Over my career, I moved from rule-based systems to classical ML, then to deep learning and not to agents. Many approaches that I used several years ago aren’t actual anymore. For example, I was using TensorFlow 1.x in 2017, then switched to PyTorch in 2019, and now I use agents which could write code in PyTorch or Jax depending on the need.
What carries over is the experience: understanding the problem, preparing data, validating models, and inspecting errors. Treat what you are learning now as your way into the first job rather than a permanent toolkit. After you get that job, learning new tools gets easier, and while models and approaches keep changing, your experience compounds.
Keep up with news, but don’t try to read about every new shiny thing
I browse Hugging Face Papers and subscribe to a couple of newsletters, such as Hacker Newsletter and Deep Learning Weekly. That doesn’t mean I read every paper or try every release. My personal assumption is that important things will come up repeatedly, so missing the first announcement is usually fine. I pay closer attention when a development relates to something I am doing. More than that, it is simply impossible to keep up with everything, so don’t try.
Read engineering writeups on corporate blogs. Look for the original problem, why the team chose its approach, and which constraints made alternatives unsuitable.
Getting a job requires separate skills

Interviews often require skills that you don’t automatically acquire by building projects or by working. You can have a great portfolio and a lot of experience and still fail interviews.
ML interviews usually include coding problems, ML theory, ML system design, and behavioral questions. I described what these rounds looked like during my job search in 2024 and how I prepared for each of them in a separate post. Both posts describe interviews for senior roles, and junior level roles usually are less strict.
Coding rounds are usually straightforward - they require you to solve LeetCode-style problems in a limited time. ML theory rounds can be more open-ended, and system design rounds often require you to discuss trade-offs and justify your choices. Behavioral questions are about your past experiences, how you work in a team, and how you handle challenges. If this is an interview for a junior role, most likely the behavioral and the design rounds will be omitted.
For coding rounds, practice on LeetCode and similar platforms. For ML theory, review the algorithms, evaluation approaches, practical aspects of working with models. For system design, think about how you would design a simple ML system, including data collection, model training, evaluation, and deployment. Behavioral interviews usually expect answers in the STAR format (Situation, Task, Action, Result). Don’t forget to prepare examples of mistakes and of things not going well.
Passing interviews is one side of the interview process. The other side is finding the right opportunity. Make sure to keep up your CV updated and your work public. Don’t just post on social media trivial things, but it is important to share your accomplishments and what you can do.
If you are changing careers, use what you already know. My economics and ERP consulting background gave me experience discussing requirements and understanding business processes. Someone who knows logistics or accounting may notice a data problem that a programmer unfamiliar with that domain would miss.
When considering a first job, ask who will review your work and how you will get feedback. In some of my early, smaller companies, I had tasks but little feedback and guidance. A larger company may offer more established support, but it depends more on the specific team.
A starting plan and a way to review progress
If you are already taking an introductory course, here is what I would do next:
- Complete the current course and its exercises, assuming it still fits your level and goal.
- Revisit the ML basics you struggled with, using small experiments to check your understanding.
- Practice programming and problem-solving, including DSA.
- Choose a small project and complete it, reaching a version that you could share.
These can overlap; you don’t have to study theory for months before starting the project. Keep the idea small enough to implement.
Once a month, write a short progress report for yourself. Record what you can now do, what you have learned, what mistakes you fixed, and your next goals.
I’ll finish this blogpost in a similar way to my previous post:
All of this is only a beginning. Following any roadmap will help you start your journey in your desired career. The rest is up to you!
blogpost career datascience machinelearning learning