torchrecurrent.benchmarks.timit#

torchrecurrent.benchmarks.timit(features, targets, batch_size=32, shuffle=True, *, feature_size=120, num_classes=180, padding_value=0.0, target_padding_value=-100, generator=None, **dataloader_kwargs)[source]#

Prepare aligned TIMIT features for frame-level phone-state recognition.

The default contract follows the TIMIT experiment of Le et al. (2015), Section 4.4 (https://arxiv.org/abs/1504.00941): each frame contains 40 log-Mel filterbank coefficients, their deltas, and accelerations (120 features), and is aligned to one of 180 Kaldi phone states.

This function does not download the licensed TIMIT corpus or reproduce the external Kaldi alignment pipeline. It batches user-provided aligned feature and target tensors for direct use with batch-first recurrent layers.

Parameters:
  • features – Utterance tensors shaped (frames, feature_size).

  • targets – Corresponding frame-label tensors shaped (frames,).

  • batch_size – Number of utterances per batch.

  • shuffle – Whether to shuffle utterances between epochs.

  • feature_size – Expected number of acoustic features per frame.

  • num_classes – Number of valid phone-state target classes.

  • padding_value – Value used to pad acoustic feature sequences.

  • target_padding_value – Value used to pad target sequences. The default matches torch.nn.CrossEntropyLoss’s ignore_index.

  • generator – Optional generator used for reproducible shuffling.

  • **dataloader_kwargs – Additional arguments passed to torch.utils.data.DataLoader.

Returns:

A data loader yielding TIMITBatch objects. Features have shape (batch, max_frames, feature_size), targets have shape (batch, max_frames), and lengths has shape (batch,).