Skip to content

LSTM

1. 🎯Objective: Live Up To The Name

When learning LSTM, I once wondered why we have so many gates, and how to interpret the so-called “cell candidate”.

Then one day, my mind realized how to interpret many of LSTM designs.

For every sequential step tt, we have to track two terms of memories:

  • Long — denoted as ctc_t. Cell state.
  • Short — denoted as hth_t. Hidden state.

All the following designs are centered at such an objective.

2. 🚪Gates

  • it=σ(xt⋅Wx(i)+ht−1⋅Wh(i)+b(i))i_t = \sigma(x_t \cdot W_x^{(i)} + h_{t - 1} \cdot W_h^{(i)} + b^{(i)})

  • ft=σ(xt⋅Wx(f)+ht−1⋅Wh(f)+b(f))f_t = \sigma(x_t \cdot W_x^{(f)} + h_{t - 1} \cdot W_h^{(f)} + b^{(f)})

  • ot=σ(xt⋅Wx(o)+ht−1⋅Wh(o)+b(o))o_t = \sigma(x_t \cdot W_x^{(o)} + h_{t - 1} \cdot W_h^{(o)} + b^{(o)})

Gates, according to word definition — represent the proportion of openness.

Proportion is of course between 0 and 1. Sigmoid function has a range of [0,1][0, 1]. That’s why 3 gates all have Sigmoid as the final stage.

Let’s see what do these gates work on.

3. 🅾️ Output Gate — Short Memory

🌳 Short Inevitably Becomes Long

Current step’s short memory hth_t can also be long memory — if you are standing at future steps.

In other words, hth_t should directly or indirectly influence ct+kc_{t + k} for k≥1k \ge 1.

Thus, hth_t is an intermediate output for the development of ct+kc_{t + k}.

But Where Does hth_t Come From 🤔

Why not let current step’s long memory ctc_t deduce hth_t 😉

Of course, in case ctc_t goes too extreme, we better normalize it into a smaller range.

[−1,1][-1, 1] is a quite small range. What function naturally has this range? Tanh.

However, Tanh only does normalization. It tells nothing about extent.

After all, we need extent to know how much should ctc_t impact hth_t.

ⓢ Sigmoid Is Here Again

Speaking of extent, its idea is quite like that of proportion — represented by Sigmoid.

This is how we end up with ht=ot⊙Tanh(ct)h_t = o_t \odot \text{Tanh}(c_t), as oto_t contains Sigmoid.

4. 🎡 Forget & Input Gates — Long Memory

How To Build Up ctc_t ⁇

ctc_t is current step’s long memory, so when establishing ctc_t, we might wanna ask:

What kind of things might be long enough for current step tt?

  • ct−1c_{t - 1} is already considered long w.r.t. step t−1t - 1, so it definitely is long for step tt. Memories can’t grow younger.

  • But……ht−1h_{t - 1} should have a chance, too. It is short at previous step t−1t - 1. However, short grows older into long.

🪘Ingredients Set

OK, so now we can mainly classify ctc_t’s ingredients into two contents:

  • Absolutely long — ct−1c_{t - 1}.

  • Relatively young among the long — ht−1h_{t - 1}. It just got promoted into long.

But there is one more content to explain.

🌊 xtx_t Shows Up As Content

Here comes a tricky part: for LSTM to learn from each step’s input xtx_t, we must let xtx_t participate in calculations of long & short memories.

With ht=ot⊙Tanh(ct)h_t = o_t \odot \text{Tanh}(c_t), xtx_t can indirectly contribute to hth_t by being part of ctc_t calculation. So the 3rd3^{rd} content is xtx_t.

Because it and ht−1h_{t - 1} are much younger than ct−1c_{t - 1}, we combine xtx_t and ht−1h_{t - 1} into cell candidate, denoted as ct~\tilde{c_t}:

ct~=Tanh(xt⋅Wx(c)+ht−1⋅Wh(c)+b(c))\tilde{c_t} = \text{Tanh}(x_t \cdot W_x^{(c)} + h_{t - 1} \cdot W_h^{(c)} + b^{(c)})

Why called “candidate”? ct~\tilde{c_t} is a candidate to compete/collaborate with ct−1c_{t - 1} for determining ctc_t.

📊Learnable Weighted Combinations

Should ct~\tilde{c_t} and ct−1c_{t - 1} compete or collaborate with each other for determining ctc_t?

This question is ought to be learnable. Naivest way is to consider 4 most extreme scenarios:

CasesDiscard ct~\tilde{c_t}Trust ct~\tilde{c_t} 100%
Discard ct−1c_{t - 1}Reset.ct~\tilde{c_t} wins competition.
Trust ct−1c_{t - 1} 100%ct−1c_{t - 1} wins competition.Collaboration.
  • Reset — ctc_t believes neither ct~\tilde{c_t} nor ct−1c_{t - 1} is helpful.

  • Collaboration — ctc_t believes both ct~\tilde{c_t} and ct−1c_{t - 1} are 100% helpful.

  • Correct~this is not zero-sum game. Remember, both sides can contribute equally.

Not hard to tell that, we should allow learning anywhere between reset and collaboration.

✚Back To Proportion Managements — Sigmoid

Now we can write ct=ft⊙ct−1+it⊙ct~c_t = f_t \odot c_{t - 1} + i_t \odot \tilde{c_t}.

ftf_t and iti_t are gates for two sides, respectively.

Let’s talk about the gate for ct−1c_{t - 1} first:

  • Notice that it is called “forget”, because we want to know how much to forget about long memory at previous step t−1t - 1.

  • Although such a naming is quite……counter-intuitive. Higher ftf_t means less forgetting, according to ctc_t formula.

  • I would have called it the “keep” gate 😏

Turn to the gate for ct~\tilde{c_t}. It is named “input”:

  • Since ct~\tilde{c_t} has the input xtx_t at current step tt.

  • Also, ct~\tilde{c_t} contains ht−1h_{t - 1}, which is previous step t−1t - 1 short memory.

  • ☝️That just promotes to an input for long memory.

5. 🎬 Into Production

As of now, it’s not hard to tell that, if we have current step’s value/label yty_t,

then we can just use hth_t to make a prediction called yt^\hat{y_t}:

yt^=f(ht⋅Wh(y)+bh(y))\hat{y_t} = \text{f}(h_t \cdot W_h^{(y)} + b_h^{(y)}) where f\text{f} is an activation function.

Some LSTM applications:

  • Time series: yty_t is each timestep’s value or label. Can be regression or classification.

  • Word/sentence generation: yty_t is each slot’s token. A token can represent a character, or even a word.

For word/sentence generation, it’s required to left-shift input, since only {yt−k  ∀  1≤k≤t}\{ y_{t - k} \; \forall \; 1 \leq k \leq t \} can be part of xtx_t to determine yty_t.

Rule of thumb: ensure everything inside xtx_t happens earlier than yty_t.