Datasets 5.0.0

5.0.0
Added 8
  • Parse Agent traces messages for SFT using the teich library to enable training on traces from claude_code, pi, codex and other sources
  • Use multiple input shards for shuffle buffer in streaming mode with configurable max_buffer_input_shards parameter
  • Add batch(by_column=...) method to batch examples by a specified column, useful for robotics datasets
  • Add Apache Iceberg format support
  • Add TsFile (Apache IoTDB) packaged builder with per-device wide format
  • Add 3D mesh support and MeshFolder builder
  • Add .conll / .conllu dataset format loader for CoNLL-2003, 2000, and U formats
  • Add num_proc argument to Dataset.to_sql
Changed 3
  • Default shuffling mechanism now uses multiple input shards instead of single shard
  • Support fsspec 2026.4.0
  • Pass library_name and version to HfApi in dataset push and delete paths
Fixed 6
  • Fix storage_options lookup for streaming Lance datasets
  • Fix Parquet streaming hangs at the end of script
  • Fix parquet reshard
  • Fix parquet columns argument
  • Fix progress bar exceeding total when load_from_cache_file=False in map operation
  • Fix single lance file from pylance 7.0
Datasets Features
Agent traces
  • Parse Agent traces messages for SFT using teich by @lhoestq in https://github.com/huggingface/datasets/pull/8232

    • Agent traces from claude_code/pi/codex and others can now be loaded with load_dataset
    • Using the teich library (new optional dependency), traces are parsed to messages to enable training on traces using e.g. trl
    • Load the data:
    >>> from datasets import load_dataset
    >>> ds = load_dataset("lhoestq/agent-traces-example", split="train")
    >>> ds[0]["messages"]
    [{'role': 'user', 'content': 'Download a random dataset from Hugging Face, use DuckDB to inspect it, and come back with a short report about it. Be concise and include: dataset name, what files/format you found, row count or rough size if you can determine it,...'
     ...]
    
    • Train on agent traces:
    trl sft --dataset-name lhoestq/agent-traces-example ...
    
Next-level shuffling in streaming mode
  • Use multiple input shards for shuffle buffer by @lhoestq in https://github.com/huggingface/datasets/pull/8194

    ds = load_dataset(..., streaming=True)
    ds = ds.shuffle(seed=42)
    # or configure local buffer shuffling manually, default is:
    ds = ds.shuffle(seed=42, buffer_size=1000, max_buffer_input_shards=10)
    

    before👎:

    after✨:

    toy example comparison

    from datasets import IterableDataset
    
    ds = IterableDataset.from_dict({"i": range(123_456_789)}, num_shards=1024)
    ds = ds.shuffle(seed=42)
    
    print("Cold start ids:")
    print(list(ds.take(10)["i"]))
    print("Nominal regime ids:")
    print(list(ds.skip(10_000).take(10)["i"]))
    

    before👎:

    Cold start ids:
    [6148853, 6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858]
    Nominal regime ids:
    [6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858, 6149290]
    

    after✨:

    Cold start ids:
    [7836668, 9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871]
    Nominal regime ids:
    [9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871, 16758448]
    

    Note: ds.state_dict() and ds.load_state_dict() are still supported for this improved shuffling :) enabling dataset checkpointing

    Note 2: it uses threads to fetch the first examples in parallel from the input shards

    Note 3: This is a BREAKING CHANGE: the default shuffling mechanism now uses multiple input shards. You can get the old mechanism by passing max_buffer_input_shards=1 to IterableDataset.shuffle()

New batching features for robotics datasets
  • Add batch(by_column=...) by @lhoestq in https://github.com/huggingface/datasets/pull/8172

    from datasets import Dataset
    
    ds = Dataset.from_dict({"episode": [0] * 10 + [1] * 10, "frame": list(range(10)) * 2})
    # ds = ds.to_iterable_dataset()
    ds = ds.batch(by_column="episode")
    for x in ds:
        print(x)
    # {'episode': [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]}
    # {'episode': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]}
    
New supported formats
Other improvements and bug fixes
New Contributors

Full Changelog: https://github.com/huggingface/datasets/compare/4.8.5...5.0.0

View original

Upgraded? How did it go?

Discussion