A way to load only needed model weights into VRAM, then swap them out.
Model paging is like cooking on a counter the size of a pizza box. One tray comes out. Another gets shoved back in the fridge.
It helps big models fit in small VRAM. You meet it in local AI apps and inference servers.
Offloading
Model paging is a finer way to move weights in and out.
VRAM
It trades time for space when VRAM is too small.
Inference engine
The inference engine chooses which page to load next.
Unified memory
Unified memory can move pages between main memory and VRAM.