All notes

The AI Stack You Should Learn Right Now

If you're a software engineer trying to break into AI, don't start by becoming a researcher. Learn the stack that companies are actually building with.

8 min read

There has probably never been a better time to enter the AI space.

And if you're a software engineer trying to move into AI, you don't necessarily need to spend the next two years learning classical machine learning from scratch.

The opportunity is increasingly around building with models, deploying them, evaluating them, and turning them into reliable products.

The stack worth understanding looks something like this:

LLMs
 ↓
Embeddings
 ↓
RAG
 ↓
Vector Databases
 ↓
Reranking
 ↓
Agents
 ↓
MCP
 ↓
Evals
 ↓
Fine-tuning
 ↓
Inference
 ↓
Quantization
 ↓
vLLM
 ↓
GPU Infrastructure

You don't need to master everything at once.

But you should know what these things are, why they exist, and where they fit.

Start with LLM fundamentals

Before jumping into agents and frameworks, understand what you're actually building with.

Learn:

The goal isn't to become an ML researcher.

The goal is to be able to answer:

"What actually happens between me sending a prompt and the model generating the next token?"

without hand-waving.

Learn Transformers

You don't need to implement a transformer from scratch on day one.

But you should understand the architecture that powers modern LLMs.

Learn:

The Hugging Face LLM Course is a good way to get familiar with the actual model ecosystem.

This is where AI starts becoming more than just:

client → API → response

You start understanding the machinery underneath.

RAG

If you're coming from backend engineering, RAG is one of the most useful areas to learn.

The basic pipeline:

Documents
   ↓
Chunking
   ↓
Embeddings
   ↓
Vector Database
   ↓
Retrieval
   ↓
Reranking
   ↓
Context
   ↓
LLM

But don't stop at building a PDF chatbot.

Learn:

Then build something real.

For example:

An AI engineering documentation agent that understands an entire codebase and answers questions with citations.

That's a much stronger project than another basic chatbot.

Agents

This is where things get interesting.

An LLM that only generates text is one thing.

An LLM that can reason, use tools, observe results, and take actions is something else.

A simplified agent loop looks like:

User
 ↓
LLM
 ↓
Reason / Plan
 ↓
Choose Tool
 ↓
Execute Tool
 ↓
Observe Result
 ↓
LLM
 ↓
Next Action

Learn frameworks such as:

But don't become framework-dependent.

Understand the underlying concepts first:

The Hugging Face Agents Course is a useful resource here.

Then build something substantial.

For example:

An AI software engineer that can inspect a GitHub repository, understand the codebase, search documentation, create an implementation plan, modify code, run tests, analyze failures, and create a PR.

That teaches you significantly more than another CRUD application.

MCP

Learn Model Context Protocol.

Think of MCP as a standardized way for AI applications to interact with external tools and data.

Understand:

Then build something.

For example:

AI Agent
   ↓
MCP Server
   ↓
 ┌───────────┬───────────┐
GitHub      Jira       Database

Build an MCP server for your own projects.

Let an AI assistant interact with your GitHub repositories, issues, documentation, databases, or other tools.

That gives you practical experience with the tool ecosystem instead of simply knowing what MCP stands for.

Evals

This is the part a lot of people skip.

Don't.

A production AI application isn't:

String response = llm.generate(prompt);

It's closer to:

Prompt
  ↓
Model
  ↓
Response
  ↓
Evaluation
  ↓
Quality
Correctness
Hallucination
Latency
Cost

Learn:

The difference between a cool AI demo and a production AI system is often evaluation and reliability.

Fine-tuning

Once you understand prompting, RAG, agents, and evaluation, learn fine-tuning.

Understand:

And most importantly, understand when not to fine-tune.

A simple mental model:

Need external / changing knowledge?
        ↓
       RAG
 
Need the model to behave differently?
        ↓
    Fine-tuning
 
Need the model to perform actions?
        ↓
 Tool Calling / Agents

Knowing this distinction is more useful than blindly fine-tuning a model because you can.

LLM inference

This is where you start moving toward AI infrastructure.

Learn what happens when a trained model actually serves requests.

Understand:

Then learn vLLM.

vLLM is particularly interesting because it takes you from:

"I know how to call an LLM."

to:

"I understand how LLM inference is actually served at scale."

Quantization

Large models are expensive.

Quantization is one of the techniques used to reduce the memory and computational requirements of models.

At a high level:

Higher precision
      ↓
More memory
More compute
Higher cost
 
Lower precision
      ↓
Less memory
Less compute
Lower cost

Learn the basics of:

You don't need to become a CUDA expert immediately.

Just understand what is being traded away when precision is reduced.

GPU Basics

If you're serious about AI engineering, learn some GPU fundamentals.

You don't need to become a hardware engineer.

Understand:

Eventually, you should be able to look at a model and reason about questions like:

How much GPU memory does this model need?

Why is inference slow?

Why does increasing batch size change throughput?

Why does quantization help?

That's where software engineering starts intersecting with AI infrastructure.

Where LangChain and LangGraph fit

Learn them.

But don't make the mistake of thinking:

"Knowing LangChain means I know AI."

Frameworks change.

The underlying concepts don't.

You should understand the architecture first and then use frameworks to move faster.

For example:

LLM
 ↓
Tool Calling
 ↓
State
 ↓
Memory
 ↓
Agent Loop
 ↓
Evaluation

LangChain and LangGraph are tools for implementing these ideas.

They're not substitutes for understanding them.

The stack to know

If you're trying to become an AI-focused software engineer, this is a useful mental model:

                    AI ENGINEERING
                          │
          ┌───────────────┴───────────────┐
          │                               │
        MODELS                         SYSTEMS
          │                               │
    LLMs / VLMs                       RAG
    Transformers                      Agents
    Fine-tuning                       MCP
    Embeddings                        Evals
    Quantization                      Observability
          │                               │
          └───────────────┬───────────────┘
                          │
                    INFRASTRUCTURE
                          │
                 vLLM / GPUs / CUDA
                 Inference / Routing
                 Caching / Latency
                 Cost Optimization

You don't need to learn all of this simultaneously.

A good progression is:

LLM Fundamentals
        ↓
Transformers
        ↓
Embeddings
        ↓
RAG
        ↓
Vector DBs
        ↓
Reranking
        ↓
Agents
        ↓
LangGraph
        ↓
MCP
        ↓
Evals
        ↓
Fine-tuning
        ↓
Inference
        ↓
Quantization
        ↓
vLLM
        ↓
GPU Basics

Don't become a certificate collector

You don't need 15 AI certificates.

A much better split is:

30% learning + 70% building.

For every major topic:

Learn
 ↓
Build
 ↓
Break
 ↓
Debug
 ↓
Read docs / papers
 ↓
Build something harder

Your GitHub should eventually contain projects that demonstrate the progression:

ai-learning/
│
├── llm-experiments/
├── embeddings/
├── rag-engine/
├── rag-evaluation/
├── agent-system/
├── langgraph-agent/
├── mcp-server/
├── ai-coding-agent/
├── fine-tuning/
├── quantization/
└── llm-inference/

That's a much stronger signal than:

"Completed 12 AI certificates."

If you're trying to get an AI job

This is probably the most important part.

You don't need to wait until you've mastered every area before applying.

Build projects while learning.

Contribute to AI-related open source.

Read AI infrastructure repositories.

Deploy models.

Build agents.

Build RAG systems.

Write about what you learn.

Contribute PRs.

Talk to engineers building these systems.

The goal isn't to become someone who knows AI terminology.

The goal is to become someone who can sit down, take an ambiguous AI problem, and build a working system around it.

And if you're wondering whether this is a good time to enter the AI space:

Yes, the opportunity is clearly there.

But don't interpret that as "learn ChatGPT APIs and put AI on your resume."

The bar is moving toward engineers who understand both software engineering and the AI stack underneath the product.

If you already know backend engineering, distributed systems, APIs, databases, caching, queues, and production systems, you have a useful foundation.

Now add:

LLMs + RAG + Agents + MCP + Evals + Fine-tuning + Inference + Quantization + vLLM + GPU basics.

That's where things get interesting.


The roadmap above is based on the uploaded material, including its emphasis on moving from LLM fundamentals through RAG, agents, MCP, evals, fine-tuning, inference, and AI infrastructure. :contentReference[oaicite:0]{index=0}