[home] [coding projects] [research projects]
Character-level RNN/LSTM modeling, context length, hidden size, and text generation.
This project studies character-level recurrent language modeling on Tiny Shakespeare. It compares a vanilla RNN with an LSTM, then separately tests the effect of context length and LSTM hidden-state size.
The first 300,000 characters are lowercased and converted to character IDs. A training example is a length-\(T\) character window and the target is the same window shifted by one character, so the model learns next-character prediction at every position.
class ShakespeareDataset(Dataset):
def __init__(
self,
encoded_text,
sequence_length=DEFAULT_SEQUENCE_LENGTH
):
self.encoded_text = encoded_text
self.sequence_length = sequence_length
def __getitem__(self, idx):
current_characters = self.encoded_text[
idx:idx+self.sequence_length
]
next_characters = self.encoded_text[
idx+1:idx+self.sequence_length+1
]
return current_characters, next_characters
Both models use a character embedding, one recurrent layer, and a linear vocabulary projection. The training loop therefore stays unchanged while the recurrent cell changes.
class VanillaRNN(nn.Module):
def __init__(self, vocab_size, embedding_dim=64,
hidden_size=64, num_layers=1):
super().__init__()
self.embedding = nn.Embedding(vocab_size, embedding_dim)
self.rnn = nn.RNN(
embedding_dim, hidden_size,
num_layers=num_layers, batch_first=True
)
self.output_layer = nn.Linear(hidden_size, vocab_size)
class LSTMModel(nn.Module):
def __init__(self, vocab_size, embedding_dim=64,
hidden_size=64, num_layers=1):
super().__init__()
self.embedding = nn.Embedding(vocab_size, embedding_dim)
self.lstm = nn.LSTM(
embedding_dim, hidden_size,
num_layers=num_layers, batch_first=True
)
self.output_layer = nn.Linear(hidden_size, vocab_size)
Generation samples from the softmax distribution one character at a time and carries the recurrent hidden state forward instead of resetting it between generated characters. The experiments compare RNN versus LSTM, sequence lengths 25 versus 50, and LSTM hidden sizes 32 versus 64.
The report finds the LSTM consistently stronger than the vanilla RNN: it converges faster, reaches lower validation perplexity, and produces more coherent generated text. Increasing hidden size has a large effect; increasing context length helps the vanilla RNN more modestly.





These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.