Part I · Foundations
Chapter 04
Running Models Locally
Self-hosting is the decision most often made for the wrong reason and abandoned for the right one — this chapter ends with the arithmetic that tells you which side of the break-even you are on.
Deliverable: A tiered Ollama + FastAPI serving stack, plus the local-versus-hosted utilisation sweep.
What's inside
11 topics
- 4.1Ollama Overview: Local LLM Runtime Engine
- 4.2Running Ollama Models with Docker
- 4.3Configuring OpenWebUI with an Ollama Backend
- 4.4FastAPI Environment Setup and Dependencies
- 4.5Integrating Ollama with FastAPI
- 4.6Configuring and Securing a Hugging Face Account
- 4.7Accessing Instruct-Tuned Models
- 4.8Installing and Using Hugging Face CLI Tools
- 4.9Model Downloading and Execution from the Hub
- 4.10Quantization, GGUF and Hardware Sizing
- 4.11Decision Guide: Local Weights versus Hosted APIs
Preparing PDF viewer…
A note on this content
The book and its chapters are my personal learning notes — compiled from online research and hands-on practice, with most of the content AI-generated from that research and learning. It is not a peer-reviewed publication, and I make no claim that it is 100% error-free. If you spot a mistake, I'd genuinely appreciate hearing about it — contact me.