[home]   [coding projects]   [research projects]


Recurrent Neural Networks

Character-level RNN/LSTM modeling, context length, hidden size, and text generation.

[full source]

This project studies character-level recurrent language modeling on Tiny Shakespeare. It compares a vanilla RNN with an LSTM, then separately tests the effect of context length and LSTM hidden-state size.

1. Sequence construction

The first 300,000 characters are lowercased and converted to character IDs. A training example is a length-\(T\) character window and the target is the same window shifted by one character, so the model learns next-character prediction at every position.

Christopher_Housholder_HW4.py
class ShakespeareDataset(Dataset):
    def __init__(
        self,
        encoded_text,
        sequence_length=DEFAULT_SEQUENCE_LENGTH
    ):
        self.encoded_text = encoded_text
        self.sequence_length = sequence_length

    def __getitem__(self, idx):
        current_characters = self.encoded_text[
            idx:idx+self.sequence_length
        ]
        next_characters = self.encoded_text[
            idx+1:idx+self.sequence_length+1
        ]
        return current_characters, next_characters

2. RNN and LSTM under the same interface

Both models use a character embedding, one recurrent layer, and a linear vocabulary projection. The training loop therefore stays unchanged while the recurrent cell changes.

Christopher_Housholder_HW4.py
class VanillaRNN(nn.Module):
    def __init__(self, vocab_size, embedding_dim=64,
                 hidden_size=64, num_layers=1):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, embedding_dim)
        self.rnn = nn.RNN(
            embedding_dim, hidden_size,
            num_layers=num_layers, batch_first=True
        )
        self.output_layer = nn.Linear(hidden_size, vocab_size)

class LSTMModel(nn.Module):
    def __init__(self, vocab_size, embedding_dim=64,
                 hidden_size=64, num_layers=1):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, embedding_dim)
        self.lstm = nn.LSTM(
            embedding_dim, hidden_size,
            num_layers=num_layers, batch_first=True
        )
        self.output_layer = nn.Linear(hidden_size, vocab_size)

3. Generation and experiments

Generation samples from the softmax distribution one character at a time and carries the recurrent hidden state forward instead of resetting it between generated characters. The experiments compare RNN versus LSTM, sequence lengths 25 versus 50, and LSTM hidden sizes 32 versus 64.

The report finds the LSTM consistently stronger than the vanilla RNN: it converges faster, reaches lower validation perplexity, and produces more coherent generated text. Increasing hidden size has a large effect; increasing context length helps the vanilla RNN more modestly.

RNN versus LSTM training loss.
RNN versus LSTM training loss.
RNN versus LSTM validation perplexity.
RNN versus LSTM validation perplexity.
Effect of sequence length on vanilla RNN training.
Effect of sequence length on vanilla RNN training.
LSTM hidden-size training comparison.
LSTM hidden-size training comparison.
LSTM hidden-size validation perplexity.
LSTM hidden-size validation perplexity.

These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.