The Impact of Transformer Architectures on Machine Translation
Transformer architectures displaced recurrent models as the dominant paradigm in neural machine translation within three years of their introduction. This paper traces that displacement, arguing that the decisive advantage was not raw accuracy but parallelisability during training, which converted translation quality into a problem of available compute. It evaluates the empirical record across high- and low-resource language pairs and considers where the architecture's assumptions still break down.
What's inside
- 1 Abstract
- 2 The Recurrent Baseline and Its Ceiling
- 3 Self-Attention as an Architectural Commitment
- 4 Empirical Gains Across Language Pairs
- 5 The Low-Resource Problem
- 6 Compute, Scale and Diminishing Returns
- 7 Conclusion
- 8 References
Unlock the full paper
Get the complete 724-word document β permanently. One-time cost, no subscription.
Cost
1,200 tokens
New here? Create an account and get 2,000 tokens free.