torchrecurrent.benchmarks.penn_treebank#
- torchrecurrent.benchmarks.penn_treebank(train_file, validation_file, test_file, batch_size=20, sequence_length=35, *, max_vocab_size=10000, eos_token='<eos>', unk_token='<unk>', device=None)[source]#
Prepare the standard word-level Penn Treebank language-model benchmark.
This adapter expects the three whitespace-tokenized text files from the Mikolov-preprocessed Penn Treebank corpus. It appends
<eos>at each newline, builds a training-only vocabulary, maps unseen validation and test words to<unk>, and creates contiguous streams for truncated BPTT. The default batch size and unroll length follow Zaremba et al. (2014) (https://arxiv.org/abs/1409.2329).Penn Treebank data is not downloaded or redistributed by this package. The original corpus is described by Marcus et al. (1993) (https://aclanthology.org/J93-2004/).
- Parameters:
train_file – Path to the preprocessed training split.
validation_file – Path to the preprocessed validation split.
test_file – Path to the preprocessed test split.
batch_size – Number of independent contiguous token streams per batch.
sequence_length – Maximum truncated-BPTT unroll length.
max_vocab_size – Maximum number of tokens, including
unk_token.eos_token – Token appended at the end of every input line.
unk_token – Token used for words outside the training vocabulary.
device – Device on which to store encoded token tensors.
- Returns:
PennTreebankCorpuscontaining train, validation, and test data loaders plus the token-to-index vocabulary. Each loader yields(inputs, targets)tensors shaped(batch_size, time_steps). The targets are the inputs shifted forward by one token.