AI Rookies

Flash Attention

Fact

A faster Attention method with lower VRAM use.

In Plain Words

Flash Attention is like washing dishes while you cook. The sink stays empty, and dinner lands sooner.

It speeds up Attention and uses less VRAM. You meet it during training and inference. It helps with long text.

Related Concepts

Attention
Flash Attention makes Attention faster and uses less VRAM.

Transformer
Transformers rely on Attention, so Flash Attention often speeds them up.

VRAM
Flash Attention uses less temporary memory and eases VRAM pressure.

Inference engine
Many inference engines build in Flash Attention to run LLMs faster.