LSTM
1. 🎯Objective: Live Up To The Name
When learning LSTM, I once wondered why we have so many gates, and how to interpret the so-called “cell candidate”.
Then one day, my mind realized how to interpret many of LSTM designs.
For every sequential step , we have to track two terms of memories:
- Long — denoted as . Cell state.
- Short — denoted as . Hidden state.
All the following designs are centered at such an objective.
2. 🚪Gates
Gates, according to word definition — represent the proportion of openness.
Proportion is of course between 0 and 1. Sigmoid function has a range of . That’s why 3 gates all have Sigmoid as the final stage.
Let’s see what do these gates work on.
3. 🅾️ Output Gate — Short Memory
🌳 Short Inevitably Becomes Long
Current step’s short memory can also be long memory — if you are standing at future steps.
In other words, should directly or indirectly influence for .
Thus, is an intermediate output for the development of .
But Where Does Come From 🤔
Why not let current step’s long memory deduce 😉
Of course, in case goes too extreme, we better normalize it into a smaller range.
is a quite small range. What function naturally has this range? Tanh.
However, Tanh only does normalization. It tells nothing about extent.
After all, we need extent to know how much should impact .
ⓢ Sigmoid Is Here Again
Speaking of extent, its idea is quite like that of proportion — represented by Sigmoid.
This is how we end up with , as contains Sigmoid.
4. 🎡 Forget & Input Gates — Long Memory
How To Build Up ⁇
is current step’s long memory, so when establishing , we might wanna ask:
What kind of things might be long enough for current step ?
is already considered long w.r.t. step , so it definitely is long for step . Memories can’t grow younger.
But…… should have a chance, too. It is short at previous step . However, short grows older into long.
🪘Ingredients Set
OK, so now we can mainly classify ’s ingredients into two contents:
Absolutely long — .
Relatively young among the long — . It just got promoted into long.
But there is one more content to explain.
🌊 Shows Up As Content
Here comes a tricky part: for LSTM to learn from each step’s input , we must let participate in calculations of long & short memories.
With , can indirectly contribute to by being part of calculation. So the content is .
Because it and are much younger than , we combine and into cell candidate, denoted as :
Why called “candidate”? is a candidate to compete/collaborate with for determining .
📊Learnable Weighted Combinations
Should and compete or collaborate with each other for determining ?
This question is ought to be learnable. Naivest way is to consider 4 most extreme scenarios:
| Cases | Discard | Trust 100% |
|---|---|---|
| Discard | Reset. | wins competition. |
| Trust 100% | wins competition. | Collaboration. |
Reset — believes neither nor is helpful.
Collaboration — believes both and are 100% helpful.
Correct~this is not zero-sum game. Remember, both sides can contribute equally.
Not hard to tell that, we should allow learning anywhere between reset and collaboration.
✚Back To Proportion Managements — Sigmoid
Now we can write .
and are gates for two sides, respectively.
Let’s talk about the gate for first:
Notice that it is called “forget”, because we want to know how much to forget about long memory at previous step .
Although such a naming is quite……counter-intuitive. Higher means less forgetting, according to formula.
I would have called it the “keep” gate 😏
Turn to the gate for . It is named “input”:
Since has the input at current step .
Also, contains , which is previous step short memory.
☝️That just promotes to an input for long memory.
5. 🎬 Into Production
As of now, it’s not hard to tell that, if we have current step’s value/label ,
then we can just use to make a prediction called :
where is an activation function.
Some LSTM applications:
Time series: is each timestep’s value or label. Can be regression or classification.
Word/sentence generation: is each slot’s token. A token can represent a character, or even a word.
For word/sentence generation, it’s required to left-shift input, since only can be part of to determine .
Rule of thumb: ensure everything inside happens earlier than .