Run your own
inference service.

Build a model API, benchmark it, and run it on a GPU. Learn how to choose an inference engine and manage performance and cost.

I’ve taught thousands of engineers and engineering leaders at Interview Kickstart.

4.9/5 · IK teaching rating

Your instructor’s background

What we’ll build
PromptText input
ModelRun locally
APIInference as a service
PlatformOperate & scale
Start locally, then build out the service.
Build on your own laptopNo GPU needed to startSmall, readable Python

What you’ll learn

Run and serve a model

Follow a prompt through tokenization and generation, then build an API around the model.

Measure performance

Benchmark latency, throughput, and memory use as you increase the load.

Choose GPUs and engines

Check GPU memory requirements and compare Transformers, vLLM, and Ollama for your workload.

Deploy and operate it

Run on rented hardware, manage access and cost, and troubleshoot slow or failed requests.

What we’ll cover

Labs 1–4 are ready; lessons 5–10 are in development. The final syllabus will be available before enrollment.

1Local inferenceRun your first model

What actually happens between a prompt and an answer?

Download a small open model. See the difference between weights on disk, a model in memory, and a running process. Follow text through tokenization, generation, and decoding.

After this lesson

Run a model locally and explain what each step does.

Python · PyTorch · TransformersRead the sample lesson
2HTTP & servicesGive the model an API

How can another program use a model without loading it?

Turn the same inference code into a long-running HTTP service. Send JSON, inspect a response, test invalid requests, and find the service’s raw metrics.

After this lesson

Send a request to your model over HTTP.

FastAPI · Uvicorn · HTTP
3Clients & networkingConnect clients and devices

What changes when the request comes from somewhere else?

Call one server from curl, Python, and JavaScript. Observe request metrics, then optionally connect a second device on the same trusted network.

After this lesson

Call your service from another program or device.

curl · Python · JavaScript
4Private accessControl access to your service

How do you decide who can use your service?

Use a private network to reach your service from another connection. Add an introductory API key and compare missing, incorrect, and valid credentials.

After this lesson

Connect to your service privately and require an API key.

Tailscale · API keys
5Latency, throughput & capacityBenchmark your model and servicePlanned

Do you actually need a GPU?

Build a CPU baseline: startup time, request latency, output tokens per second, and memory use. Compare the first request with steady-state requests. Change prompt length, output length, and concurrent demand, then use the results to decide whether a GPU would help.

After this lesson

Run a repeatable benchmark and explain what the results tell you about capacity.

Benchmarking · Load testing · CPU vs GPU
6Queues & observabilityHandle more requestsPlanned

What happens when requests arrive faster than you can serve them?

Explore bounded queues, admission limits, streaming, deadlines, and batching. Separate time spent waiting from time spent computing.

After this lesson

Recognize overload and choose how your service should handle it.

Scheduling · Metrics · Streaming
7Runtime decisionsChoose an inference enginePlanned

What should your application own, and what should the engine do?

Compare a plain Transformers process with vLLM and Ollama. Look at GPU memory, batching, request scheduling, and API compatibility. Define a workload and measurements for the GPU experiment in the next lesson.

After this lesson

Choose an engine for your workload and know what to test.

Transformers · vLLM · Ollama
8GPU deployment & benchmarkingRun your service on a GPUPlanned

How do you move a working local process onto another machine?

Package the service in Docker, then deploy to a Linux machine with an NVIDIA GPU. Check GPU memory, drivers, and runtime compatibility. Run vLLM, compare latency and throughput with your CPU baseline, calculate compute cost, and shut down the rented resources.

After this lesson

Deploy on a compatible GPU and explain the performance and cost you measured.

Docker · NVIDIA GPUs · vLLM
9Production boundariesOperate an inference APIPlanned

What does an API need beyond a successful response?

Extend the introductory key into identity, permissions, key rotation, quotas, TLS, and useful operational signals. Learn which responsibilities stay with the operator.

After this lesson

Manage access, usage limits, and monitoring for your service.

Identity · Quotas · Observability
10Reliability & economicsScale and maintain the servicePlanned

Who keeps the service available when demand and machines change?

Explore replicas, load balancing, cold starts, scaling, rollouts, and recovery. Compare self-operated and managed inference by responsibilities, behavior, and cost.

After this lesson

Plan for changes in traffic and recover when a server fails.

Replicas · Autoscaling · Recovery

How classes work

Follow along

I’ll walk through the code and explain each step.

Try it yourself

Run the lab and change the code.

Ask questions

We’ll work through what’s confusing or stuck.

Read a sample lesson

Support beyond the lessons

Direct access to me

Get candid feedback on your code, technical decisions, and career questions during the cohort.

Interview preparation

Talk through technical questions and get feedback on how you explain your work.

Meet other engineers

Compare approaches, work through labs together, and stay connected after the course.

Introductions and referrals

Where there’s a fit, I can help through my network. Referrals depend on the role and relationship and don’t guarantee an interview or job.

A longer-term plan

I’d like to bring in more instructors for mock interviews, feedback, and referrals. For now, I teach and support this cohort myself.

Samwel Emmanuel

Samwel Emmanuel

AI Inference & Post Training @ Baseten

I work on AI inference and post-training at Baseten. Before that, I worked at Databricks, Google, and Salesforce. I studied electrical engineering at Harvard.

Work, teaching & education

  • Education
  • Now
  • Previously
  • Previously
  • Teaching
  • Previously

Over about three years at Interview Kickstart, I’ve taught thousands of engineers and engineering leaders, including classes on AI agents and voice systems.

4.9/5 · IK teaching rating
My LinkedIn profile

An independent course by Samwel Emmanuel.

Learner feedback

Quotes from past IK classes will appear here with learners’ permission.

Who this course is for

For developers who want to run and serve models themselves. No prior inference experience needed.

A good starting point

  • You can read a small Python program.
  • You can work in a terminal.
  • You have a laptop for the labs.

Useful to know

  • Current labs are tested on macOS.
  • Some network experiments use a second device.
  • Later GPU labs may involve separate compute costs.

Questions about the course

Who is this course for?

Developers comfortable with Python and a terminal. You don’t need prior experience serving models.

Do I need a GPU or a cloud account?

The first four labs run a small open model on your laptop’s CPU. Later lessons are planned to use rented GPUs. I’ll share any expected compute costs before enrollment opens.

Can I follow along on Windows or Linux?

The walkthroughs are tested on macOS. The concepts apply to Windows and Linux, but their setup commands aren’t verified yet. Some network labs use a second device and a hotspot.

What if I miss a live session?

The cohort will include recordings. I’ll share the schedule and how long you’ll have access to the recordings before enrollment opens.

When does the first cohort start?

I’m still setting the dates. Tuition for the founding cohort is $1,499. I’ll share the schedule and weekly time commitment before enrollment opens.

Is all of the curriculum available now?

Labs 1–4 are ready; lessons 5–10 are in development. You’ll see the final syllabus before enrolling.

What kind of career support is included?

Interview preparation, feedback on your work, and introductions or referrals where there’s a fit. Referrals don’t guarantee interviews or jobs. A wider instructor community and dedicated mock interviews are future plans.

Is this an independent course?

Yes. The organizations in my bio don’t sponsor or endorse this course.

Try a lesson before you decide.

Try the sample lesson