torchrecurrent.benchmarks.penn_treebank#

torchrecurrent.benchmarks.penn_treebank(train_file, validation_file, test_file, batch_size=20, sequence_length=35, *, max_vocab_size=10000, eos_token='<eos>', unk_token='<unk>', device=None)[source]#

Prepare the standard word-level Penn Treebank language-model benchmark.

This adapter expects the three whitespace-tokenized text files from the Mikolov-preprocessed Penn Treebank corpus. It appends <eos> at each newline, builds a training-only vocabulary, maps unseen validation and test words to <unk>, and creates contiguous streams for truncated BPTT. The default batch size and unroll length follow Zaremba et al. (2014) (https://arxiv.org/abs/1409.2329).

Penn Treebank data is not downloaded or redistributed by this package. The original corpus is described by Marcus et al. (1993) (https://aclanthology.org/J93-2004/).

Parameters:
  • train_file – Path to the preprocessed training split.

  • validation_file – Path to the preprocessed validation split.

  • test_file – Path to the preprocessed test split.

  • batch_size – Number of independent contiguous token streams per batch.

  • sequence_length – Maximum truncated-BPTT unroll length.

  • max_vocab_size – Maximum number of tokens, including unk_token.

  • eos_token – Token appended at the end of every input line.

  • unk_token – Token used for words outside the training vocabulary.

  • device – Device on which to store encoded token tensors.

Returns:

PennTreebankCorpus containing train, validation, and test data loaders plus the token-to-index vocabulary. Each loader yields (inputs, targets) tensors shaped (batch_size, time_steps). The targets are the inputs shifted forward by one token.