ponytail

AIMIT

ponytail release notes.

Latest v4.9.0 · by ponytailWritten in JavaScriptWebsiteDietrichGebert/ponytailRSS

Release activity

Release activity — 11 releases across 7 days since Jun 14, 2026. Each cell is one day; darker means more releases that day. Nothing is recorded before Jun 14, 2026. Older weeks are hidden at this screen width.
JunJulAugSep
Sunday1 release on Jun 14, 2026No releases on Jun 21, 2026No releases on Jun 28, 2026No releases on Jul 5, 2026No releases on Jul 12, 2026No releases on Jul 19, 2026No releases on Jul 26, 2026No releases on Aug 2, 2026No releases on Aug 9, 2026No releases on Aug 16, 2026No releases on Aug 23, 2026No releases on Aug 30, 2026No releases on Sep 6, 2026
Monday3 releases on Jun 15, 2026No releases on Jun 22, 20261 release on Jun 29, 2026No releases on Jul 6, 2026No releases on Jul 13, 2026No releases on Jul 20, 2026No releases on Jul 27, 2026No releases on Aug 3, 2026No releases on Aug 10, 2026No releases on Aug 17, 2026No releases on Aug 24, 2026No releases on Aug 31, 2026No releases on Sep 7, 2026
Tuesday1 release on Jun 16, 20262 releases on Jun 23, 2026No releases on Jun 30, 2026No releases on Jul 7, 2026No releases on Jul 14, 2026No releases on Jul 21, 2026No releases on Jul 28, 2026No releases on Aug 4, 2026No releases on Aug 11, 2026No releases on Aug 18, 2026No releases on Aug 25, 2026No releases on Sep 1, 2026No releases on Sep 8, 2026
WednesdayNo releases on Jun 17, 20262 releases on Jun 24, 2026No releases on Jul 1, 2026No releases on Jul 8, 2026No releases on Jul 15, 2026No releases on Jul 22, 2026No releases on Jul 29, 2026No releases on Aug 5, 2026No releases on Aug 12, 2026No releases on Aug 19, 2026No releases on Aug 26, 2026No releases on Sep 2, 2026
ThursdayNo releases on Jun 18, 2026No releases on Jun 25, 2026No releases on Jul 2, 2026No releases on Jul 9, 2026No releases on Jul 16, 2026No releases on Jul 23, 2026No releases on Jul 30, 2026No releases on Aug 6, 2026No releases on Aug 13, 2026No releases on Aug 20, 2026No releases on Aug 27, 2026No releases on Sep 3, 2026
FridayNo releases on Jun 19, 2026No releases on Jun 26, 2026No releases on Jul 3, 2026No releases on Jul 10, 2026No releases on Jul 17, 2026No releases on Jul 24, 2026No releases on Jul 31, 20261 release on Aug 7, 2026No releases on Aug 14, 2026No releases on Aug 21, 2026No releases on Aug 28, 2026No releases on Sep 4, 2026
SaturdayNo releases on Jun 20, 2026No releases on Jun 27, 2026No releases on Jul 4, 2026No releases on Jul 11, 2026No releases on Jul 18, 2026No releases on Jul 25, 2026No releases on Aug 1, 2026No releases on Aug 8, 2026No releases on Aug 15, 2026No releases on Aug 22, 2026No releases on Aug 29, 2026No releases on Sep 5, 2026

11 releases since Jun 14, 2026, busiest day 3

Changelog

v4.9.0Latest

v4.9.0: 53 commits of doing less

Added 4
  • Add /ponytail default <mode> command to persist the default mode across restarts
  • Add Qoder hooks support (UserPromptSubmit + PreToolUse)
  • Add opt-in agent-type scoping for SubagentStart injection via PONYTAIL_SUBAGENT_MATCHER
  • Add pi status bar controls to hide the indicator while keeping ponytail active
Changed 2
  • Make bare /ponytail command report the active level instead of resetting to default
  • Make the ultra statusline badge stand out
Fixed 14
  • Fix copilot-plugin test to check all 6 command files
  • Fix exec lifecycle hook commands
  • Handle stdin error in ponytail-mode-tracker to avoid uncaught crash
  • Prevent Windows session freeze from stdin EOF deadlock
  • Strip UTF-8 BOM before parsing config.json
  • Don't destroy combined statuslines on uninstall

From ponytail

Five weeks of merged work in one drop. The headline: /ponytail default <mode> makes your chosen mode survive restarts, and a bare /ponytail now reports the active level instead of resetting it. Also new: Qoder support, subagent scoping via PONYTAIL_SUBAGENT_MATCHER, pi status bar controls, and about 30 fixes across Windows, Codex, OpenCode, and the uninstaller.

What's Changed
New Contributors

Full Changelog: https://github.com/DietrichGebert/ponytail/compare/v4.8.4...v4.9.0

View originalPermalink
How v4.9.0 went
v4.8.4

v4.8.4: lazy in Hermes now

Added 2
  • Ponytail now runs as a native Hermes Agent plugin with always-on context, bundled skills, and slash commands
  • Ponytail ships as a Devin CLI plugin
Changed 1
  • The skill now triggers on any coding task, not just keyword prompts, improving recall on coding tasks from 2/6 to 6/6 with precision unchanged
Fixed 2
  • Benchmarks now strip block comments before counting lines of code
  • Benchmark rejects --ollama-url without a host

From ponytail

New
  • Ponytail now runs as a native Hermes Agent plugin (#323) — always-on context, bundled skills, and slash commands.
  • Ponytail ships as a Devin CLI plugin (#318).
  • The skill now triggers on any coding task, not just keyword prompts (#447) — description-driven hosts load it reliably on plain coding work, recall on coding tasks went 2/6 to 6/6 with precision unchanged.
Docs
  • Sponsors section, with GreenPT as the first sponsor (#321)
  • Announcement banner replaces the badge; waitlist teaser added and translated (#316, #314, #312)
  • Swift/SwiftUI section in platform-native.md (#313)
  • Star History chart and Trendshift badges in the READMEs (#294, #293)
  • ES and Korean READMEs synced with the Devin CLI addition (#322)
Fixes
  • benchmarks: strip block comments before counting LOC (#232)
  • benchmark: reject --ollama-url without a host (#315)

Full changelog: https://github.com/DietrichGebert/ponytail/compare/v4.8.3...v4.8.4

View originalPermalink
How v4.8.4 went
v4.8.3

v4.8.3: lazy in subagents too

Added 1
  • The ponytail ruleset now injects into subagents via a SubagentStart hook, ensuring spawned subagents stay lazy like the main thread

From ponytail

New
  • The ponytail ruleset now injects into subagents via a SubagentStart hook (#254) — spawned subagents stay lazy, not just the main thread.
Docs
  • Korean README — README.ko.md (#283)
  • OpenCode npm install + npm badge across the English and Spanish READMEs; dropped the obsolete command-symlink note (#285)

Full changelog: https://github.com/DietrichGebert/ponytail/compare/v4.8.2...v4.8.3

View originalPermalink
How v4.8.3 went
v4.8.2

v4.8.2: now on npm

Added 4
  • Package is now installable from npm as @dietrichgebert/ponytail
  • Pi-extension status bar indicator for the active ponytail mode
  • Swival install instructions
  • PowerShell-safe shared hooks
Fixed 2
  • Uninstall cleanup for state left outside plugin files
  • Benchmark and test fixes

From ponytail

Ponytail is now installable from npm as @dietrichgebert/ponytail, published from CI via npm trusted publishing (OIDC, with provenance) — no tokens.

OpenCode
{ "plugin": ["@dietrichgebert/ponytail"] }
Also in this release
  • pi-extension status bar indicator for the active ponytail mode (#275, #279)
  • Swival install instructions (#264)
  • PowerShell-safe shared hooks (#265)
  • uninstall cleanup for state left outside plugin files (#228)
  • benchmark + test fixes (#274, #268)

Full changelog: https://github.com/DietrichGebert/ponytail/compare/v4.8.1...v4.8.2

View originalPermalink
How v4.8.2 went
v4.8.1

v4.8.1: consistent versioning

Added 1
  • Add a CI guard (scripts/check-versions.js) that fails when version files disagree, or when a release tag does not match the manifest version
Fixed 1
  • Align all four plugin manifests and both package.json files to 4.8.1

From ponytail

Patch release. Metadata only, no changes to the ruleset or behavior.

v4.8.0 shipped with the version manifests still reading 4.7.0 (and the package.json files at 0.1.0), so plugins reported 4.7.0 as the latest version even after updating. This release aligns every version-bearing file to 4.8.1 so the version reports correctly across Claude, Codex, Copilot, and Gemini.

Fixes
  • Align all four plugin manifests and both package.json files to 4.8.1 (#260, #262)
  • Add a CI guard (scripts/check-versions.js) that fails when version files disagree, or when a release tag does not match the manifest version, so this cannot recur
Updating

Update the ponytail plugin in your host (re-run your marketplace or plugin update, or reinstall). It will then report 4.8.1.

View originalPermalink
How v4.8.1 went
v4.8.0

v4.8.0: comprehension first, now with an MCP server

Added 5
  • Add ponytail-mcp, an MCP server that serves the ruleset to any MCP-capable agent
  • Add /ponytail-gain measured-impact scoreboard showing less code, less cost, and more speed from benchmark medians
  • Add Antigravity support via .agents/rules/ponytail.md
  • Add CodeWhale support with native AGENTS.md reader and zero setup
  • Add argument-hint on the ponytail skill
Fixed 12
  • Implement comprehension-first guard and reuse rung to climb the ladder only after understanding the problem and reusing what already lives in the repo
  • Do not embed shell-unsafe install paths in the statusline setup nudge
  • Make Python command in robustness-audit.js portable for Windows
  • Do not write output on SessionStart for Copilot
  • Register skills dir via config hook so OpenCode discovers ponytail skills
  • Avoid Gemini loading Claude hook events
Security 1
  • Bump @modelcontextprotocol/sdk to ^1.26.0 to address CVE-2026-25536

From ponytail

He read the whole thing first, then wrote less.

This release sharpens the core ruleset (comprehension before laziness, reuse before rewriting), ships an MCP server for the rules, adds a measured-impact scoreboard, and lands a security bump. Two more agents join the family.

✨ Features
  • ponytail-mcp: an MCP server that serves the ruleset to any MCP-capable agent (#91)
  • /ponytail-gain: measured-impact scoreboard — less code, less cost, more speed from the benchmark medians (#108)
  • Antigravity support via .agents/rules/ponytail.md (#119)
  • CodeWhale support — native AGENTS.md reader, zero setup (#124)
  • argument-hint on the ponytail skill (#85)
🔒 Security
  • Bump @modelcontextprotocol/sdk to ^1.26.0 (CVE-2026-25536) (#208)
🛠 Fixes
  • Comprehension-first guard + reuse rung — climb the ladder only after understanding the problem; reuse what already lives in the repo before rewriting it (#245, #217)
  • Don't embed shell-unsafe install paths in the statusline setup nudge (#224)
  • Portable Python command in robustness-audit.js (Windows) (#209)
  • Don't write output on SessionStart for Copilot (#168)
  • Register skills dir via config hook so OpenCode discovers ponytail skills (#138)
  • Avoid Gemini loading Claude hook events (#139)
  • Statusline reads the flag from CLAUDE_CONFIG_DIR, not just ~/.claude (#154)
  • Guard final writeHookOutput against stdout EPIPE in ponytail-activate (#149)
  • Strip UTF-8 BOM before parsing settings.json in ponytail-activate (#148)
  • Resolve test failure on Node.js < 20.11.0 (#157)
  • Only deactivate on a standalone "stop ponytail" / "normal mode" (#162)
  • Pin all four safety carve-outs in the rule-drift canary (#114)
📊 Benchmarks
  • Agentic LOC + safety benchmark against a fair agentic baseline (#126, #158)
  • Completeness judge so LOC wins can't hide under-delivery (#171)
  • critic-email task reproducing the critique's own example (#173)
📚 Docs & examples
  • Spanish (LATAM) README translation (#110)
  • 6 new over-engineering survivors + platform-native guide (#109)
  • Agentic benchmark chart and corrected cost claim (42–75%, 30-rep) in the README (#160, #129)

Full changelog: https://github.com/DietrichGebert/ponytail/compare/v4.7.0...v4.8.0

View originalPermalink
How v4.8.0 went
v4.7.0

v4.7.0: lazy in OpenClaw now

Added 2
  • Add OpenClaw skill that integrates ponytail as an agent capability accessible via clawhub install ponytail
  • Include review, audit, debt, and help skills as part of the OpenClaw integration
Changed 1
  • Generate the OpenClaw skill directly from ponytail's single source to prevent drift across platforms

From ponytail

OpenClaw is the fastest-growing open-source agent out there, an always-on assistant that reads your messages and runs your workflows. The lazy senior dev now lives inside it.

clawhub install ponytail and he's there: an OpenClaw skill that kicks in on coding tasks and tells the agent to write less. The review, audit, debt, and help skills come along too. He does not care how big the house is. He still deletes more than he adds.

The skill is generated straight from ponytail's single source, so the OpenClaw copy cannot drift from every other platform. One ruleset, now on one more agent.

Tested the boring way: installed OpenClaw, loaded the skill, watched it come up ready and visible to the model. Then he went back to not talking.

View originalPermalink
How v4.7.0 went
v4.6.0

v4.6.0: help, reluctantly

Added 3
  • Add `/ponytail-help` command that lists all available commands, wired up on all skill-capable hosts (Claude Code, Codex, OpenCode, Gemini CLI, pi)
  • Add Ollama runner to test ponytail on local models
  • Add parity test to ensure no future command can be advertised without corresponding implementation files
Fixed 2
  • Fix counter that scored unfenced code as zero
  • Fix Unicode character that crashed the runner on Windows

From ponytail

He has never explained a command in his life. You typed /ponytail-help and got nothing, because the file was never actually there, only the promise of it in the docs. The most senior-dev bug there is: works in the standup, missing from the repo.

Now it ships. /ponytail-help is wired up alongside the other commands on every skill-capable host (Claude Code, Codex, OpenCode, Gemini CLI, pi): one command that lists the rest. A new parity test makes sure no future command can be advertised without the files to back it. He hates writing documentation. He hates broken promises more.

We benchmarked him on a tiny local model and shipped the flop.

A contributor added an Ollama runner so you can test ponytail on local models. We ran it on llama3.2 (3B) and the lines-of-code win turned out to be noise: one run lands 17% under baseline, the next 50% over, the median shrugs. The skill is tuned for models that actually follow instructions. A 3B model nods along and writes the boilerplate anyway.

We published that instead of burying it. A benchmark you only show when it flatters you is an ad. The frontier numbers (80-94% less code on Haiku, Sonnet, Opus) still hold, and now there's an honest note on where they stop. Full reproduction in benchmarks/results/.

Swept up on the way out: a counter that scored unfenced code as zero, and a Unicode character that crashed the runner on Windows after the work was already done.

He'd call it a quiet release. Then he'd stop talking.

View originalPermalink
How v4.6.0 went
v4.5.0

v4.5.0: lazy in Copilot

Added 2
  • GitHub Copilot CLI plugin support via Copilot Marketplace
  • CI on every push and PR with npm test script
Fixed 2
  • Correctness checks now use python3 for macOS/CI compatibility
  • Hooks degrade gracefully when node is not on PATH

From ponytail

This release is mostly other people, and that's the point. GitHub Copilot CLI is now a full plugin host, contributed by @maxfelker (a Microsoft engineer) who built it with Copilot, tested it live, then reviewed his own PR in ponytail ultra (the headline finding was "delete a test"). CI, an npm test script, and a python3 fix came from @christophermayfield. Plus a fix for the hooks erroring when node isn't on PATH (Nix/nvm setups). The tool that does less got more thorough by getting more hands.

  • GitHub Copilot CLI plugin: copilot plugin marketplace add DietrichGebert/ponytail then copilot plugin install ponytail@ponytail.
  • CI on every push and PR, plus npm test.
  • Correctness checks work on macOS/CI now (python3 probe).
  • Hooks degrade gracefully when node isn't on PATH.
What's Changed
New Contributors

Full Changelog: https://github.com/DietrichGebert/ponytail/compare/v4.4.0...v4.5.0

View originalPermalink
How v4.5.0 went
v4.4.0

v4.4.0: field-tested, still lazy

Added 3
  • Add ponytail-debt skill to harvest deferred ponytail: shortcuts into a ledger so later doesn't quietly become never
  • Add a behavior-gate eval so rules can't silently regress
  • Add a dark-background logo
Changed 2
  • Sharpen rules based on field review feedback: hardware is never the spec ideal (leave the calibration knob), the one-runnable-check rule is now a headline (lazy code without its check is unfinished), and explanation you explicitly asked for isn't debt
  • Use the dark logo in the README header on dark themes

From ponytail

The headline of this release isn't a feature, it's a field test. A user ran ponytail across a from-scratch rewrite of a real system: nine phases, protocol plus desktop app plus simulator plus Raspberry Pi daemon plus ESP32 firmware. The verdict was "net win, kept it on the whole build," and across all nine phases "it never once trimmed a failsafe, validation, or auth check." It also flagged where the laziness needed a tighter leash. v4.4.0 is the result.

  • Sharper rules from that feedback: hardware is never the spec ideal (leave the calibration knob), the one-runnable-check rule is now a headline ("lazy code without its check is unfinished"), and explanation you explicitly asked for isn't debt.
  • /ponytail-debt: harvests the ponytail: shortcuts you've deferred into a ledger, so "later" doesn't quietly become "never."
  • A behavior-gate eval so those rules can't silently regress.
  • A dark-background logo (community-contributed) and a cleaner README.
What's Changed

Full Changelog: https://github.com/DietrichGebert/ponytail/compare/v4.3.0...v4.4.0

View originalPermalink
How v4.4.0 went
v4.3.0

v4.3.0: more agents, still lazy

Added 3
  • Add ponytail-audit skill
  • Add Gemini CLI support
  • Add correctness assertion to benchmarks
Fixed 2
  • Fix Windows hooks failing under PowerShell when cmd.exe %VAR% is not expanded
  • Honor CLAUDE_CONFIG_DIR in hooks

From ponytail

What's Changed
New Contributors

Full Changelog: https://github.com/DietrichGebert/ponytail/compare/v4.2.0...v4.3.0

View originalPermalink
How v4.3.0 went
View all

Discussion

If you publish ponytail, you can claim this product by proving you administer its repository.