Skip to content
EntityQ85810444· pop 37· linked from 1,242 articles

transformer

Sign in to save

Also known as transformer model, transformer architecture, transformers

machine-learning model architecture first developed by Google Brain

Wikidata facts

Show 2 more facts
time of discovery or invention
2017-06-12
Sources (2)

via Wikidata · CC0

~40 min read

Article

A standard transformer architecture. Many modern diagrams show the pre-layer normalization (pre-LN) convention, while the original 2017 paper used post-layer normalization (post-LN).

In deep learning, the transformer is a family of artificial neural network architectures built around the attention mechanism. Transformers were introduced to model sequential data without recurrence and without convolutions, allowing much more parallel computation during training. They are now a dominant architecture for natural language processing, computer vision, speech processing, multimodal learning, robotics, and many other sequence-modelling tasks.

Connections

Categories