# ultralytics v8.4.133 — v8.4.133 - Improve hyperparameter Tuner mutation convergence (#25984) - Product: ultralytics (https://whatsnew.fyi/product/ultralytics) - Vendor: ultralytics - Date: 2026-08-29 - Version: v8.4.133 - Original notes: https://github.com/ultralytics/ultralytics/releases/tag/v8.4.133 - Permalink: https://whatsnew.fyi/product/ultralytics/releases/v8.4.133 What's New is an index, not a publisher: every entry below links to the vendor's own release notes, which are the authoritative source. Entries are labelled where they are hand-curated sample data, pre-releases, or drawn from a secondary source such as a developer blog. Reuse: the summaries, labels and curation here are © What's New. Quote freely with attribution and a link back; wholesale republication of the corpus is not permitted — terms: https://whatsnew.fyi/terms. The vendors' own release notes remain their publishers'. --- - **changed** — Replace coordinate-by-coordinate crossover with fitness-weighted selection of complete, high-performing configurations in hyperparameter tuning - **changed** — Mutate approximately half of the parameters in normalized search-space coordinates to allow parameters starting at zero to evolve more effectively - **changed** — Gradually reduce mutation size when tuning stops finding better results to encourage refinement after broad exploration - **fixed** — Prevent duplicate candidates after clipping, rounding, or integer conversion in hyperparameter tuning - **changed** — Set Ray Tune to default to Optuna multivariate TPE with parallel-aware suggestions instead of independent random search - **changed** — Move image channel reordering and tensor-contiguity operations from CPU-side NumPy processing to the inference device - **added** — Enable channels-last memory layout automatically for native PyTorch inference and standalone validation on supported x86 Linux and Windows CPUs with oneDNN - **fixed** — Fix fraction handling during classification and detection INT8 export calibration so scalar fractions apply directly to the selected calibration split - **added** — Custom detection datasets can now report small-, medium-, and large-object mAP when using save_json=True - **changed** — Simplify edge-device installation by installing the base ultralytics package instead of the larger [export] extra, with export dependencies installed automatically when an export is requested - **changed** — W&B model artifact uploads now follow the existing training save argument where save=False skips uploading the best checkpoint while retaining metrics and plots - **changed** — Convert saved models back to safe contiguous format and clear stale EMA data to improve compatibility with channels-last inference ##### 🌟 Summary **Ultralytics 8.4.133 improves hyperparameter tuning convergence, speeds up inference preprocessing, expands detection metrics, and simplifies edge-device setup.** 🚀 ##### 📊 Key Changes - **Smarter hyperparameter tuning — PR #25984 by @glenn-jocher** - Replaces coordinate-by-coordinate crossover with **fitness-weighted selection of complete, high-performing configurations**. - Preserves useful relationships between hyperparameters instead of mixing them independently. - Mutates approximately half of the parameters in normalized search-space coordinates, allowing parameters that start at zero—such as `degrees` or `shear`—to evolve more effectively. - Gradually reduces mutation size when tuning stops finding better results, encouraging refinement after broad exploration. - Prevents duplicate candidates after clipping, rounding, or integer conversion, including small and discrete search spaces. - Ray Tune now defaults to **Optuna multivariate TPE**, with parallel-aware suggestions rather than independent random search. - **Faster predictor preprocessing — PR #25982 by @jahsef** ⚡ - Moves image channel reordering and tensor-contiguity operations from CPU-side NumPy processing to the inference device. - Preserves output values while reducing unnecessary CPU copies. - Reported benchmarks show approximately **2.2–3.1× faster preprocessing on an RTX 5080**, with additional gains on CPU. - **Automatic channels-last CPU inference — PR #25983 by @JESUSROYETH** - Enables channels-last memory layout automatically for native PyTorch inference and standalone validation on supported x86 Linux and Windows CPUs with oneDNN. - Keeps training defaults and unsupported platforms unchanged. - Explicit `channels_last=True` remains available for supported CPU and CUDA paths. - Saved models are converted back to a safe contiguous format and stale EMA data is cleared to improve compatibility. - **More accurate INT8 calibration subsets — PR #25978 by @JESUSROYETH** - Fixes `fraction` handling during classification and detection INT8 export calibration. - Scalar fractions now apply directly to the selected calibration split, while list-based fractions retain train/validation/test behavior. - Prevents exports from unintentionally calibrating on an entire dataset when only a subset was requested. - **Size-specific mAP for custom detection datasets — PR #25981 by @fcakyon** 📈 - Custom detection datasets can now report small-, medium-, and large-object mAP when using `save_json=True`. - Builds temporary COCO-format annotations internally while preserving existing native metrics and prediction files. - Applies consistently during training validation, final-model validation, and standalone validation. - **Simpler edge-device installation** - Raspberry Pi, Jetson, DGX Spark, DeepStream, and related guides now install the base `ultralytics` package instead of the larger `[export]` extra. - Export dependencies are installed automatically when an export is requested, reducing installation size and dependency conflicts. - **Improved Weights & Biases artifact control — PR #25985 by @glenn-jocher** - W&B model artifact uploads now follow the existing training `save` argument. - `save=False` skips uploading the best checkpoint while retaining metrics and plots. - Default behavior remains unchanged with `save=True`. - **Package update** - Version bumped to **8.4.133**. ##### 🎯 Purpose & Impact - **Better tuning results:** Hyperparameter searches are more likely to preserve successful configurations, explore meaningful alternatives, and avoid wasting trials on duplicates. 🎯 - **Faster inference:** Device-side preprocessing can reduce latency, particularly for batched inference and CPU-bound pipelines. - **Broader performance optimization:** Supported x86 CPU users may benefit from channels-last inference without changing their existing commands. - **More reliable model export:** INT8 cali _[Truncated at 4000 characters — full notes: https://github.com/ultralytics/ultralytics/releases/tag/v8.4.133]_