/
Recurrent and Transformer Models Overview
Save to my account
Sign up
Recurrent and Transformer Models Overview
Recurrent and Transformer Models Overview
Study
1
Question
What are recurrent neural networks (RNNs) established as state-of-the-art approaches for?
Answer
RNNs, especially Long Short-Term Memory (LSTM) and Gated Recurrent Neural Networks (GRNN), are established as state-of-the-art approaches in sequence modeling and transduction problems such as language modeling and machine translation.
2
Question
What is a fundamental constraint of recurrent models in sequence processing?
Answer
The inherently sequential nature of recurrent models precludes parallelization within training examples, which becomes critical at longer sequence lengths due to memory constraints limiting batching across examples.
3
Question
How have recent works improved computational efficiency in recurrent models?
Answer
Recent works have achieved significant improvements in computational efficiency through factorization tricks and conditional computation, while also improving model performance in the case of the latter.
4
Question
What role do attention mechanisms play in sequence modeling?
Answer
Attention mechanisms allow modeling of dependencies without regard to their distance in the input or output sequences.
5
Question
What is the Transformer model's approach to sequence modeling?
Answer
The Transformer model architecture relies entirely on an attention mechanism to draw global dependencies between input and output, eschewing recurrence.
6
Question
What is a key advantage of the Transformer model in terms of training time?
Answer
The Transformer can reach a new state of the art in translation quality after being trained for as little as twelve hours on eight P100 GPUs.
7
Question
What are some models that have the goal of reducing sequential computation?
Answer
The Extended Neural GPU, ByteNet, and ConvS2S are models that use convolutional neural networks as basic building blocks to compute hidden representations in parallel.
8
Question
What is self-attention in the context of the Transformer architecture?
Answer
Self-attention is an attention mechanism that relates different positions of a single sequence in order to compute a representation of the sequence.
9
Question
How does the Transformer differ from traditional sequence transduction models?
Answer
The Transformer is the first transduction model that relies entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
10
Question
What is the encoder-decoder structure in competitive neural sequence transduction models?
Answer
In the encoder-decoder structure, the encoder maps an input sequence of symbol representations to a sequence of continuous representations, and the decoder generates an output sequence of symbols one element at a time.
11
Question
What components are included in the Transformer model architecture?
Answer
The Transformer model architecture includes the Encoder and Decoder, both of which consist of stacks of identical layers containing modules for Multi-Head Attention, Add & Norm, and Feed Forward Network.
12
Question
What is the output generation process in the decoder of the Transformer model?
Answer
The decoder generates an output sequence by being auto-regressive, consuming previously generated symbols as additional input when generating the next.
13
Question
What is 'Scaled Dot-Product Attention' in the Transformer?
Answer
Scaled Dot-Product Attention computes the output as a weighted sum based on the relevance of queries and keys, where the weights are derived from the dot products of the queries and keys, scaled by the square root of their dimension.
14
Question
Why is scaling used in the Scaled Dot-Product Attention?
Answer
Scaling is used to counteract the effect of large dot products pushing the softmax function into regions where it has extremely small gradients, which could hinder learning.
15
Question
What is Multi-Head Attention in the Transformer model?
Answer
Multi-Head Attention allows the model to jointly attend to information from different representation subspaces at different positions by performing multiple attention functions in parallel.
16
Question
What are positional encodings and why are they used in the Transformer?
Answer
Positional encodings are added to input embeddings to inject information about the relative or absolute position of tokens in the sequence, allowing the model to utilize the order of the sequence.
17
Question
What mathematical functions are used for positional encodings in the Transformer?
Answer
The positional encodings are computed using sine and cosine functions of different frequencies.
18
Question
What are the three desiderata considered in comparing self-attention layers to recurrent and convolutional layers?
Answer
The three desiderata are total computational complexity per layer, the amount of computation that can be parallelized, and the maximum path length between long-range dependencies.
19
Question
What is the dimensionality of the input and output in the feed-forward networks of the Transformer?
Answer
The dimensionality of input and output is d_model 512, and the inner-layer has dimensionality d_ff 2048.
20
Question
What is the purpose of layer normalization in the Transformer?
Answer
Layer normalization is used around each sub-layer to stabilize and improve the training process.
21
Question
What does the term 'auto-regressive' mean in the context of the Transformer decoder?
Answer
'Auto-regressive' means that the model generates each output symbol one at a time, using previously generated symbols as additional input for generating the next symbol.
22
Question
What is the complexity of self-attention compared to recurrent and convolutional layers?
Answer
Self-attention has a complexity of O(n^2) for d-dimensional outputs, while recurrent layers have O(n*d^2) and convolutional layers have O(k*n*d^2) complexity.
23
Question
What types of tasks has self-attention been successfully used for?
Answer
Self-attention has been used successfully in tasks including reading comprehension, abstractive summarization, textual entailment, and learning task-independent sentence representations.
24
Question
What is the significance of masking in the self-attention mechanism of the decoder?
Answer
Masking in the self-attention mechanism of the decoder prevents future position attention, preserving the auto-regressive property.
25
Question
What is the connection method used by a self-attention layer?
Answer
A self-attention layer connects all positions with a constant number of sequentially executed operations.
26
Question
How do recurrent layers connect input and output positions?
Answer
A recurrent layer requires O(n) sequential operations to connect input and output positions.
27
Question
When are self-attention layers faster than recurrent layers?
Answer
Self-attention layers are faster than recurrent layers when the sequence length n is smaller than the representation dimensionality d.
28
Question
What representations are commonly used in state-of-the-art models for machine translation?
Answer
Common representations used are word-piece and byte-pair representations.
29
Question
What is one proposed solution to improve computational performance for long sequences in self-attention?
Answer
Restrict self-attention to consider only a neighborhood of size r in the input sequence centered around the output position.
30
Question
What is the maximum path length introduced by restricting self-attention?
Answer
The maximum path length increases to O(nr).