About the paper
BBT preserves the sequence-compression benefits of byte-pair encoding (BPE) while replacing vocabulary lookup and the vocabulary-sized language model head with byte-level composition. It composes segment representations from bytes and generates each segment byte by byte.
Key results
- Lower test bits per byte than matched standard BPE Transformers
- 20–40% fewer parameters at comparable FLOPs
- Improved robustness to character-level corruption
- Stronger transfer to unseen languages