Inference/Lesson 1/Free preview

What happens when
you load a model?

In this sample from the first lesson, we’ll look at how a Python program loads a model and uses it to answer prompts.

Lesson 1 · Local inferenceShort reading and practice questionNo setup required for this preview

Downloading and loading a model

When you download a model, you get files: learned weights, a configuration, and the tokenizer’s rules. They sit on disk. Downloading those files doesn’t start anything.

Your Python program loads the tokenizer and model into memory. Now you have a running process that can turn text into tokens, generate more tokens, and turn the result back into text.

From model files to a response
DISK                  RUNNING PYTHON PROCESS
model files   ──────▶  tokenizer + model in memory
                               │
                               ▼
prompt → format → tokenize → generate → decode → text

Reusing a loaded model

In the first lab, model loading happens before the prompt loop. Each new prompt repeats the text-processing and generation steps. It reuses the model that is already loaded.

Simplified teaching pseudocode
tokenizer, model = load_from_disk()

while True:
    prompt = read_prompt()
    if prompt == "quit":
        break

    tokens = tokenize(format_prompt(prompt))
    new_tokens = generate(model, tokens)
    print(decode(new_tokens))

This pseudocode shows the main steps; it won’t run as written. In the full lab, we’ll use a small open model and walk through the Python code for each step.

Check your understanding

You send a second prompt.
What needs to happen again?

The Python process is still running, and your first response has finished. Choose the best explanation.

What happens when the process stops?

When the process exits, its memory is released. The files on disk remain. Next time, you load those files again—you don’t need to download them again.

Check your understanding: Explain the difference between downloading a model, loading it, and running inference. Then explain why a model can stay loaded without remembering your last conversation.

Calling the model from another program

So far, we’ve entered prompts directly in the terminal. In the next lesson, we’ll wrap the same code in an HTTP service so another program can send prompts and receive answers.

From lesson 2 · Your first model API

Which request runs the model?

The server has already loaded the model. Read these two requests and make a prediction before you check.

GET /health
POST /generate
{"prompt": "Say hello."}
Check your answer

Only POST /generate calls the model. GET /health returns a fixed status. It tells you the app can respond, but it doesn’t test generation speed or answer quality.

In the lab, you run both requests and inspect the code and metrics to see the difference.

See what we’ll cover next.

Read the course outline and details about the first live cohort.

Explore the course