Downloading and loading a model
When you download a model, you get files: learned weights, a configuration, and the tokenizer’s rules. They sit on disk. Downloading those files doesn’t start anything.
Your Python program loads the tokenizer and model into memory. Now you have a running process that can turn text into tokens, generate more tokens, and turn the result back into text.
DISK RUNNING PYTHON PROCESS
model files ──────▶ tokenizer + model in memory
│
▼
prompt → format → tokenize → generate → decode → textReusing a loaded model
In the first lab, model loading happens before the prompt loop. Each new prompt repeats the text-processing and generation steps. It reuses the model that is already loaded.
tokenizer, model = load_from_disk()
while True:
prompt = read_prompt()
if prompt == "quit":
break
tokens = tokenize(format_prompt(prompt))
new_tokens = generate(model, tokens)
print(decode(new_tokens))This pseudocode shows the main steps; it won’t run as written. In the full lab, we’ll use a small open model and walk through the Python code for each step.
You send a second prompt.
What needs to happen again?
The Python process is still running, and your first response has finished. Choose the best explanation.
What happens when the process stops?
When the process exits, its memory is released. The files on disk remain. Next time, you load those files again—you don’t need to download them again.
Calling the model from another program
So far, we’ve entered prompts directly in the terminal. In the next lesson, we’ll wrap the same code in an HTTP service so another program can send prompts and receive answers.
From lesson 2 · Your first model API
Which request runs the model?
The server has already loaded the model. Read these two requests and make a prediction before you check.
GET /healthPOST /generate
{"prompt": "Say hello."}Check your answer
Only POST /generate calls the model. GET /health returns a fixed status. It tells you the app can respond, but it doesn’t test generation speed or answer quality.
In the lab, you run both requests and inspect the code and metrics to see the difference.
See what we’ll cover next.
Read the course outline and details about the first live cohort.
Explore the course