Run and serve a model
Follow a prompt through tokenization and generation, then build an API around the model.
Build a model API, benchmark it, and run it on a GPU. Learn how to choose an inference engine and manage performance and cost.
I’ve taught thousands of engineers and engineering leaders at Interview Kickstart.
Your instructor’s background

Follow a prompt through tokenization and generation, then build an API around the model.
Benchmark latency, throughput, and memory use as you increase the load.
Check GPU memory requirements and compare Transformers, vLLM, and Ollama for your workload.
Run on rented hardware, manage access and cost, and troubleshoot slow or failed requests.
Labs 1–4 are ready; lessons 5–10 are in development. The final syllabus will be available before enrollment.
Download a small open model. See the difference between weights on disk, a model in memory, and a running process. Follow text through tokenization, generation, and decoding.
Run a model locally and explain what each step does.
Turn the same inference code into a long-running HTTP service. Send JSON, inspect a response, test invalid requests, and find the service’s raw metrics.
Send a request to your model over HTTP.
Call one server from curl, Python, and JavaScript. Observe request metrics, then optionally connect a second device on the same trusted network.
Call your service from another program or device.
Use a private network to reach your service from another connection. Add an introductory API key and compare missing, incorrect, and valid credentials.
Connect to your service privately and require an API key.
Build a CPU baseline: startup time, request latency, output tokens per second, and memory use. Compare the first request with steady-state requests. Change prompt length, output length, and concurrent demand, then use the results to decide whether a GPU would help.
Run a repeatable benchmark and explain what the results tell you about capacity.
Explore bounded queues, admission limits, streaming, deadlines, and batching. Separate time spent waiting from time spent computing.
Recognize overload and choose how your service should handle it.
Compare a plain Transformers process with vLLM and Ollama. Look at GPU memory, batching, request scheduling, and API compatibility. Define a workload and measurements for the GPU experiment in the next lesson.
Choose an engine for your workload and know what to test.
Package the service in Docker, then deploy to a Linux machine with an NVIDIA GPU. Check GPU memory, drivers, and runtime compatibility. Run vLLM, compare latency and throughput with your CPU baseline, calculate compute cost, and shut down the rented resources.
Deploy on a compatible GPU and explain the performance and cost you measured.
Extend the introductory key into identity, permissions, key rotation, quotas, TLS, and useful operational signals. Learn which responsibilities stay with the operator.
Manage access, usage limits, and monitoring for your service.
Explore replicas, load balancing, cold starts, scaling, rollouts, and recovery. Compare self-operated and managed inference by responsibilities, behavior, and cost.
Plan for changes in traffic and recover when a server fails.
I’ll walk through the code and explain each step.
Run the lab and change the code.
We’ll work through what’s confusing or stuck.
Get candid feedback on your code, technical decisions, and career questions during the cohort.
Talk through technical questions and get feedback on how you explain your work.
Compare approaches, work through labs together, and stay connected after the course.
Where there’s a fit, I can help through my network. Referrals depend on the role and relationship and don’t guarantee an interview or job.
I’d like to bring in more instructors for mock interviews, feedback, and referrals. For now, I teach and support this cohort myself.

AI Inference & Post Training @ Baseten
I work on AI inference and post-training at Baseten. Before that, I worked at Databricks, Google, and Salesforce. I studied electrical engineering at Harvard.
Work, teaching & education

Over about three years at Interview Kickstart, I’ve taught thousands of engineers and engineering leaders, including classes on AI agents and voice systems.
My LinkedIn profileAn independent course by Samwel Emmanuel.
Quotes from past IK classes will appear here with learners’ permission.
For developers who want to run and serve models themselves. No prior inference experience needed.
Developers comfortable with Python and a terminal. You don’t need prior experience serving models.
The first four labs run a small open model on your laptop’s CPU. Later lessons are planned to use rented GPUs. I’ll share any expected compute costs before enrollment opens.
The walkthroughs are tested on macOS. The concepts apply to Windows and Linux, but their setup commands aren’t verified yet. Some network labs use a second device and a hotspot.
The cohort will include recordings. I’ll share the schedule and how long you’ll have access to the recordings before enrollment opens.
I’m still setting the dates. Tuition for the founding cohort is $1,499. I’ll share the schedule and weekly time commitment before enrollment opens.
Labs 1–4 are ready; lessons 5–10 are in development. You’ll see the final syllabus before enrolling.
Interview preparation, feedback on your work, and introductions or referrals where there’s a fit. Referrals don’t guarantee interviews or jobs. A wider instructor community and dedicated mock interviews are future plans.
Yes. The organizations in my bio don’t sponsor or endorse this course.