Inference Engineering - Part 6: Diffusion, and What Ties It Together
Every earlier part of this series assumed autoregressive generation: the model writes one token at a time, left to right, and each token depends on all the ones before...
Page 1 of 11
Every earlier part of this series assumed autoregressive generation: the model writes one token at a time, left to right, and each token depends on all the ones before...
Three parameters sit between a trained model and its output: temperature, top_k and top_p. They are widely used and rarely defined precisely. All three act at the same...
The idea here is simple, but memory bandwidth comes up so often in inference engineering that it deserves a post of its own. It rests on the split Part 2 drew between ...
Many times, you would come across the claim that “the encoder builds an internal representation of the input.” However, most modern LLMs are decoder-only. So do they j...
Part 1 ended with a few hundred gigabytes of weights and an architecture description. This post is about what happens to them next, and it is where the series title st...
Every hotel reservation system starts with the same line of code: SELECT 1 FROM reservation WHERE room_id = $1 AND check_out > $2 AND check_in < $3; -- no rows?...
I kept running into large language model (LLM) systems and finding that I could use them without understanding them. That is an uncomfortable place to be: you end up a...
Part 7 argued that a groundedness gate has to score atomic claims and take the minimum, because a five-sentence answer with one fabrication averages out to a passing s...
Seven posts of machinery, and none of it is trustworthy until you can answer one question: how do you know it works? RAG systems are unusually hostile to intuition her...
Retrieval can succeed completely and the answer can still be wrong. The model receives five relevant chunks and asserts a sixth thing that none of them support — a pla...