# Datasets: what changed from 4 to 5 - Product: Datasets (https://whatsnew.fyi/product/hugging-face-datasets) - Vendor: Hugging Face - Range: changelog entries numbered after 4.8.5 up to and including 5.0.1, stable releases only - Entries below: 2 releases (newest first) - Resolved: 4 is 4.8.5 and 5 is 5.0.1, the newest stable release of each major we track - Carrying security changes: 1 · CVEs mentioned: 0 · Mentioning breaking changes: 1 · Removing or deprecating something: 0 - Page: https://whatsnew.fyi/product/hugging-face-datasets/compare/4...5 What's New is an index, not a publisher: every entry below links to the vendor's own release notes, which are the authoritative source. Entries are labelled where they are hand-curated sample data, pre-releases, or drawn from a secondary source such as a developer blog. Reuse: the summaries, labels and curation here are © What's New. Quote freely with attribution and a link back; wholesale republication of the corpus is not permitted — terms: https://whatsnew.fyi/terms. The vendors' own release notes remain their publishers'. ## What changed (38 changes, grouped by kind) ### Added #### 5.0.1 (2026-07-28) - Support hermes traces - Support droid agent traces - Support batched=True in Dataset.to_dict #### 5.0.0 (2026-06-05) - Parse Agent traces messages for SFT using the teich library to enable training on traces from claude_code, pi, codex and other sources - Use multiple input shards for shuffle buffer in streaming mode with configurable max_buffer_input_shards parameter - Add batch(by_column=...) method to batch examples by a specified column, useful for robotics datasets - Add Apache Iceberg format support - Add TsFile (Apache IoTDB) packaged builder with per-device wide format - Add 3D mesh support and MeshFolder builder - Add .conll / .conllu dataset format loader for CoNLL-2003, 2000, and U formats - Add num_proc argument to Dataset.to_sql ### Changed #### 5.0.0 (2026-06-05) - Default shuffling mechanism now uses multiple input shards instead of single shard - Support fsspec 2026.4.0 - Pass library_name and version to HfApi in dataset push and delete paths ### Fixed #### 5.0.1 (2026-07-28) - Fix version string in __init__.py - Fix conda build - Fix JSON loader schema inference for files starting with a UTF-8 BOM - Fix batch(by_column=...) crashing after shard/shuffle/split - Fix traces streaming - Fix lance auth - Fix resuming dataloader twice resetting the dataloader - Validate Arrow IPC record batches - Make the dataset fingerprint independent of Arrow chunking - Fix casting a nullable LargeList to a different inner type - Fix column drop in Arrow path of axis=1 concatenation - Raise on length mismatch in batched IterableDataset.map - Fix require_storage_embed recursing into require_storage_cast - Fix hdf5 external files - Keep flat numeric columns with nulls numeric in numpy format - Fix CSV loader dropping on_bad_lines/encoding_errors on pandas 2.0-2.2 #### 5.0.0 (2026-06-05) - Fix storage_options lookup for streaming Lance datasets - Fix Parquet streaming hangs at the end of script - Fix parquet reshard - Fix parquet columns argument - Fix progress bar exceeding total when load_from_cache_file=False in map operation - Fix single lance file from pylance 7.0 ### Security #### 5.0.1 (2026-07-28) - Fix symlink-following arbitrary file write in archive extraction - Fix path traversal via metadata file_name in folder-based builders ## Release notes ### 5.0.1 - Date: 2026-07-28 - Version: 5.0.1 - Original notes: https://github.com/huggingface/datasets/releases/tag/5.0.1 - Permalink: https://whatsnew.fyi/product/hugging-face-datasets/releases/5.0.1 - **fixed** — Fix version string in __init__.py - **fixed** — Fix conda build - **fixed** — Fix JSON loader schema inference for files starting with a UTF-8 BOM - **added** — Support hermes traces - **fixed** — Fix batch(by_column=...) crashing after shard/shuffle/split - **fixed** — Fix traces streaming - **added** — Support droid agent traces - **fixed** — Fix lance auth - **security** — Fix symlink-following arbitrary file write in archive extraction - **fixed** — Fix resuming dataloader twice resetting the dataloader - **fixed** — Validate Arrow IPC record batches - **fixed** — Make the dataset fingerprint independent of Arrow chunking - **fixed** — Fix casting a nullable LargeList to a different inner type - **fixed** — Fix column drop in Arrow path of axis=1 concatenation - **fixed** — Raise on length mismatch in batched IterableDataset.map - **fixed** — Fix require_storage_embed recursing into require_storage_cast - **security** — Fix path traversal via metadata file_name in folder-based builders - **added** — Support batched=True in Dataset.to_dict - **fixed** — Fix hdf5 external files - **fixed** — Keep flat numeric columns with nulls numeric in numpy format - **fixed** — Fix CSV loader dropping on_bad_lines/encoding_errors on pandas 2.0-2.2 ##### Bug fixes * Fix version string in __init__.py by @qgallouedec in https://github.com/huggingface/datasets/pull/8244 * fix conda build by @lhoestq in https://github.com/huggingface/datasets/pull/8250 * Fix JSON loader schema inference for files starting with a UTF-8 BOM (#8241) by @archievi in https://github.com/huggingface/datasets/pull/8243 * Support hermes traces by @lhoestq in https://github.com/huggingface/datasets/pull/8255 * Fix batch(by_column=...) crashing after shard/shuffle/split by @pkooij in https://github.com/huggingface/datasets/pull/8259 * fix traces streaming by @lhoestq in https://github.com/huggingface/datasets/pull/8277 * support droid agent traces by @cfahlgren1 in https://github.com/huggingface/datasets/pull/8263 * Fix CI: commit operation equality (hfh 1.20.0) and pytest parametrize collection error by @Wauplin in https://github.com/huggingface/datasets/pull/8283 * Fix lance auth by @lhoestq in https://github.com/huggingface/datasets/pull/8301 * Fix symlink-following arbitrary file write in archive extraction by @AAtomical in https://github.com/huggingface/datasets/pull/8303 * Bug Fix: Resuming Twice Resets the Dataloader by @francesco-bertolotti in https://github.com/huggingface/datasets/pull/8295 * Bump fsspec and simpler wds compr by @lhoestq in https://github.com/huggingface/datasets/pull/8337 * fix: validate Arrow IPC record batches by @XciD in https://github.com/huggingface/datasets/pull/8350 * Make the dataset fingerprint independent of Arrow chunking by @SuryanshSS1011 in https://github.com/huggingface/datasets/pull/8339 * Fix casting a nullable LargeList to a different inner type by @vineethsaivs in https://github.com/huggingface/datasets/pull/8346 * docs: replace AutoFeatureExtractor with AutoImageProcessor in image preprocessing docs by @gautamkishore in https://github.com/huggingface/datasets/pull/8326 * Fix column drop in Arrow path of axis=1 concatenation by @ebarkhordar in https://github.com/huggingface/datasets/pull/8342 * Raise on length mismatch in batched IterableDataset.map by @sohumt123 in https://github.com/huggingface/datasets/pull/8332 * Fix require_storage_embed recursing into require_storage_cast by @vineethsaivs in https://github.com/huggingface/datasets/pull/8349 * Fix path traversal via metadata file_name in folder-based builders by @Kaif10 in https://github.com/huggingface/datasets/pull/8325 * Support batched=True in Dataset.to_dict by @vineethsaivs in https://github.com/huggingface/datasets/pull/8333 * Fix hdf5 external files by @lhoestq in https://github.com/huggingface/datasets/pull/8355 * remove bad require_storage test by @lhoestq in https://github.com/huggingface/datasets/pull/8357 * ensure fiels are in repo by @lhoestq in https://github.com/huggingface/datasets/pull/8356 * Keep flat numeric columns with nulls numeric in numpy format by @ebarkhordar in https://github.com/huggingface/datasets/pull/8352 * Fix bucket dataset card handling and push metadata accounting by @pjh4993 in https://github.com/huggingface/datasets/pull/8354 * Fix CSV loader dropping on_bad_lines/encoding_errors on pandas 2.0-2.2 by @ebarkhordar in https://github.com/huggingface/datasets/pull/8358 * Keep integers on the python read path for fixed-shape ArrayXD columns with nulls by @ebarkhordar in https://github.com/huggingface/datasets/pull/8363 * Rebatch arrow source before formatting in IterableDataset.filter to fix resume data loss by @ebarkhordar in https://github.com/huggingface/datasets/pull/8360 * Decode Json() columns in Dataset.to_pandas() by @ebarkhordar in https://github.com/huggingface/datasets/pull/8344 * fix buckets on windows by @lhoestq in https://github.com/huggingface/datasets/pull/8369 * Fix DatasetDict.push_to_hub leaving removed splits in the dataset card by @pjh4993 in https://github.com/huggingface/datasets/pull/8367 * Preserve nullable integer columns in to_json/to_csv/to_sql by @ebarkhordar in https://github.com/huggingface/datasets/pull _[Truncated at 4000 characters — full notes: https://github.com/huggingface/datasets/releases/tag/5.0.1]_ ### 5.0.0 - Date: 2026-06-05 - Version: 5.0.0 - Original notes: https://github.com/huggingface/datasets/releases/tag/5.0.0 - Permalink: https://whatsnew.fyi/product/hugging-face-datasets/releases/5.0.0 - **added** — Parse Agent traces messages for SFT using the teich library to enable training on traces from claude_code, pi, codex and other sources - **added** — Use multiple input shards for shuffle buffer in streaming mode with configurable max_buffer_input_shards parameter - **added** — Add batch(by_column=...) method to batch examples by a specified column, useful for robotics datasets - **added** — Add Apache Iceberg format support - **added** — Add TsFile (Apache IoTDB) packaged builder with per-device wide format - **added** — Add 3D mesh support and MeshFolder builder - **added** — Add .conll / .conllu dataset format loader for CoNLL-2003, 2000, and U formats - **added** — Add num_proc argument to Dataset.to_sql - **changed** — Default shuffling mechanism now uses multiple input shards instead of single shard - **changed** — Support fsspec 2026.4.0 - **changed** — Pass library_name and version to HfApi in dataset push and delete paths - **fixed** — Fix storage_options lookup for streaming Lance datasets - **fixed** — Fix Parquet streaming hangs at the end of script - **fixed** — Fix parquet reshard - **fixed** — Fix parquet columns argument - **fixed** — Fix progress bar exceeding total when load_from_cache_file=False in map operation - **fixed** — Fix single lance file from pylance 7.0 ##### Datasets Features ###### Agent traces * Parse Agent traces messages for SFT using `teich` by @lhoestq in https://github.com/huggingface/datasets/pull/8232 * Agent traces from claude_code/pi/codex and others can now be loaded with load_dataset * Using the `teich` library (new optional dependency), traces are parsed to `messages` to enable training on traces using e.g. `trl` * Load the data: ```python >>> from datasets import load_dataset >>> ds = load_dataset("lhoestq/agent-traces-example", split="train") >>> ds[0]["messages"] [{'role': 'user', 'content': 'Download a random dataset from Hugging Face, use DuckDB to inspect it, and come back with a short report about it. Be concise and include: dataset name, what files/format you found, row count or rough size if you can determine it,...' ...] ``` * Train on agent traces: ```bash trl sft --dataset-name lhoestq/agent-traces-example ... ``` * find all the Agent traces datasets on HF here: https://huggingface.co/datasets?format=format:agent-traces&sort=trending ###### Next-level shuffling in streaming mode * Use multiple input shards for shuffle buffer by @lhoestq in https://github.com/huggingface/datasets/pull/8194 ```python ds = load_dataset(..., streaming=True) ds = ds.shuffle(seed=42) # or configure local buffer shuffling manually, default is: ds = ds.shuffle(seed=42, buffer_size=1000, max_buffer_input_shards=10) ``` before👎: image after✨: image toy example comparison ```python from datasets import IterableDataset ds = IterableDataset.from_dict({"i": range(123_456_789)}, num_shards=1024) ds = ds.shuffle(seed=42) print("Cold start ids:") print(list(ds.take(10)["i"])) print("Nominal regime ids:") print(list(ds.skip(10_000).take(10)["i"])) ``` before👎: ``` Cold start ids: [6148853, 6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858] Nominal regime ids: [6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858, 6149290] ``` after✨: ``` Cold start ids: [7836668, 9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871] Nominal regime ids: [9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871, 16758448] ``` Note: `ds.state_dict()` and `ds.load_state_dict()` are still supported for this improved shuffling :) enabling dataset checkpointing Note 2: it uses threads to fetch the first examples in parallel from the input shards Note 3: This is a BREAKING CHANGE: the default shuffling mechanism now uses multiple input shards. You can get the old mechanism by passing `max_buffer_input_shards=1` to `IterableDataset.shuffle()` ###### New batching features for robotics datasets * Add batch(by_column=...) by @lhoestq in https://github.com/huggingface/datasets/pull/8172 ```python from datasets import Dataset ds = Dataset.from_dict({"episode": [0] * 10 + [1] * 10, "frame": list(range(10)) * 2}) # ds = ds.to_iterable_dataset() ds = ds.batch(by_column="episode") for x in ds: print(x) # {'episode': [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]} # {'episode': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]} ``` ###### New supported formats * Add Apache Iceberg format support by @frankliee in https://github.com/huggingface/datasets/pull/8148 * feat: add TsFile (Apache IoTDB) packaged builder with per-device wide format by @JackieTien97 in https://githu _[Truncated at 4000 characters — full notes: https://github.com/huggingface/datasets/releases/tag/5.0.0]_