Study notes on neural networks, Transformers, attention, decoder-only LLMs, mixture-of-experts models, and the prefill/decode split.
Posts tagged with LLM
LLM
LLM
Study notes on neural networks, Transformers, attention, decoder-only LLMs, mixture-of-experts models, and the prefill/decode split.