To me, orchestration is not really the core problem. The bigger challenge is handling agent coherency over long periods of time, (drift) so my solution is a living forum-like system where agents shout to one another, document work in items, and follow doctrine handed down by the developer and work against a known specification that's checkable. I have used this workflow in several large projects that MUST be correct, and it works for me. Currently cleaning it up for other users.
pitched 7 hours ago [-]
Not saying this counts specifically but I love seeing devs who hate Jira build something like Jira to solve managing a team of agents. It’s a beautifully ironic thing to watch.
jredwards 1 hours ago [-]
Gas Town did the same thing back in January. It's probably got three more layers of abstraction on top of it by now.
I think we'll see it more and more... open-ended Q&A / novel work is everyone's first experience of LLMs, but I don't think it's actually representative of most of the work people do. Coding kind of sits in the middle as it's novel but semi-structured.
Chat is great for novel open-ended work, but is terrible at concurrency - i.e. managing > 1 agent runs. As agents get cheaper, I suspect we'll see more more bounded processes where you want to run lots of items through them at once, and thus Jira-like things... (we ended up having this problem and building something like this, albeit more medieval themed [0])
We also found the other piece that Jira-style solves (and that chat makes super hard) is multi-player. I've yet to see a good implementation of a multi-player chat solution to agents working concurrently; and we found that an async model is the only sane solution to this.
I don't think people hate issue tracking system, they just hate Jira.
redhale 6 hours ago [-]
And I think most people who hate Jira actually just hate the process their company has imposed inside of Jira.
The vanilla out of the box product is actually quite close to old Trello, in simplicity. But there are so many ways to fuck it up and overcomplicate it, and so many companies proceed to fuck it up and overcomplicate it.
saguntum 11 minutes ago [-]
+1
I have learned recently that I don't inherently hate JIRA, I hate our JIRA and byzantine process flows.
I think because other task trackers like Trello are simpler and have less config options (per my last use, long ago, pre-Atlassian-acquisition), it's impossible knock on wood to get into the antipatterns I've seen in those tools.
Ironically I think one of the main goals of management I have seen consistently over time - ensuring tickets are accurate, up to date, and reflect reality - would end up much more true if the task tracking tools were just much simpler, with fewer fancy features.
greysonp 4 minutes ago [-]
I actually just hate Jira :) It's horrendously slow for what it does, the search is mediocre at best, and I regularly run into bugs.
greggh 6 hours ago [-]
This is absolutely it for me. I've never had to use a Jira instance that wasn't absolutely ruined by its maintainers.
Zambyte 6 hours ago [-]
Yup, I literally just hate their web UI. Once I found the Atlassian-blessed jira "cli" [0] (after searching for a long time for a tui, which is actually what the "cli" is primarily) I have actually be fine with using Jira again. As an added bonus, being a cli has made it easy to interface with using my AI tooling.
They released twg mainly for agents which is totally fine and useful as well.
alemanek 2 hours ago [-]
2nd this but I did need to write a session start hook to prune out most of the skills so it was just the non Rovo ones. Took about 15min to setup though so I am not complaining. twg works really well.
x-complexity 3 hours ago [-]
Exactly. A plain basic issue tracking system is all that's needed for 90% of cases.
Need a specific feature? Plug that in later when you actually need it.
j_bum 5 hours ago [-]
> The local-first work board for teams of coding agents. Nothing ships until a second agent verifies it.
> gate: failing
Gave me a good chuckle. Cool idea though, and I like the records. Reminds me a bit of the OpenAI hacking incidents.
Olscore 5 hours ago [-]
Haha. Locally it was passing. That badge is GitHub running the tests on its own machines, where a couple of new browser tests fail. Fix is going through the same loop: one agent fixes, another verifies, then it merges.
woah 5 hours ago [-]
> docker-agent lets you create and run intelligent AI agents that collaborate to solve complex problems — no code required.
How is "no code required" a selling point when we all have agents that can almost instantly puke out all the mostly-correct code you could want?
crossroadsguy 11 minutes ago [-]
Code means a bit of mental engagement, more than yaml, json etc. You either review that or not review that, but it's an addition. While I am not touching anything made by docker-the-bloater anytime soon, I could think that at least.
blakeashleyjr 11 hours ago [-]
While I love new open source dev (especially in Go!), agent harnesses are turning into JS frameworks from yesteryear.
All the cool kids have one!
crossroadsguy 8 minutes ago [-]
[delayed]
pjmlp 11 hours ago [-]
Coupled with VSCode forks to manage them
yipinwong 8 hours ago [-]
You know which one will win for sure. ones with the worst experience.
Vue and svelte are easy to use and learn/adopt. But look what won. React.js.
So whatever the ai agent with the vast user basae will win regardless.
---
I know it's a slippery slope, but given you brought up JS framwork analogy, I had to fall into it
ravenstine 6 hours ago [-]
It's all about what story you can tell.
Cursor is easily the worst experience I've had with an AI harness, yet it was acquired for billions of dollars in spite of being a middleman tied to a ripoff of VS Code. None of the colleagues of mine who were touting it last year are still using it. But if you can make yourself look like the next big thing, you'll get money thrown at you. Just look at Omarchy.
A good experience necessarily takes time and careful thought. Nobody these days can do that while also remaining relevant, or even surviving.
shepherdjerred 20 minutes ago [-]
Many people genuinely prefer React over Vue/Svelte
v3ss0n 1 hours ago [-]
A good model with svelte make things buttery smooth.
girvo 8 hours ago [-]
I mean as someone who was there at the time, early React was easy to learn and adopt. It’s complexity came later after it had already won
Now Angular v1, man… if I ever see another digest loop error in my life, I will scream
ravenstine 6 hours ago [-]
So glad I never had to deal with Angular, especially v1. I tried learning it and it just seemed like it would actually make applications more incomprehensible than before. React originally addressed the actual problem that most frameworks mistook for not being that important. I still don't particularly like React, but it makes sense why it became popular.
Terr_ 7 hours ago [-]
"Look, LLMs will free you from the tyranny of frameworks!"
"Uh, OK, so how do I get this thing working? Where are the docs?"
"Well, first you install its unique framework..."
aklingel 4 hours ago [-]
[flagged]
caskeycoding 10 hours ago [-]
[flagged]
verdverm 10 hours ago [-]
Go doesn't have a good mod/plugin story for harness devs to provide to their users
I say this as a gopher who also has a custom harness written in Go, it's a real challenge compared to the TS setups, but TS plugins are also a security concern
edoceo 42 minutes ago [-]
Somehow Mattermost, written in Go, allows plugins.
verdverm 22 minutes ago [-]
yup, with an http API, not a plugin the way most harnesses are doing it today
gandreani 9 hours ago [-]
I agree! Comes with the territory of a compiled language I suppose?
I haven't given this much thought (I also just started with golang last year) but I'm assuming the only real path here is to provide extension SDKs in Lua or other languages whose interpreters have been implemented in go.
I think that's what grafana's k6 [1] does with a JS interpreter
I was thinking I recalled something like that, but the nice thing about the TS ecosystem is that you can write the same language, importing and using the harness sdk types.
Go has the plugin package, but I haven't actually seen anyone use it (outside of toys and blog posts)
fails at step two, precompiled binary is already built, you can't change the code
why not just use an SDK where you are using a plugin?
otterley 7 hours ago [-]
Why isn’t IPC good enough?
verdverm 6 hours ago [-]
My custom harness has a dagger based environment, binary runs in the host, every tool call and file modification creates a new layer in a container - for better session time travel and forking
It may work for some things, but you really want to be wired into that system to do the more interesting things.
otterley 5 hours ago [-]
I don't understand why dynamically loading code into a Linux process is essential to make this work. Can you explain?
verdverm 5 hours ago [-]
I didn't say it couldn't work, it's just way more complexity and effort with IPC or Go than TS. That's one of the reasons harnesses are choosing TS. It's also a very popular language.
If, like me, you couldn't find any security-related info on the linked page.
gregwebs 9 hours ago [-]
Docker Agent is a harness. There is a sandbox mode that can be used to run it in docker sandbox (a VM, not a container). If you don't use sandbox mode then I assume it is running in a container.
If you don't want to use their harness then you wouldn't use docker agent and instead use their `sbx` cli to run the harness of your choice (claude, codex, pi, etc).
I think it will be great if Docker can get people used to using secure VMs. I am developing a similar project (still a work in progress): https://github.com/gregwebs/agent-vm
genghisjahn 9 hours ago [-]
I've found sbx to be very helpful. really like the sentinel value wrapper they have going so you can add secrets but the model can't see them. Outbound calls get looked up by sentinel value and the real one goes out to whatever API you're auth'ing too. There's other features but just having claude run in a sandbox and easily see what it does and doesn't have access to has been great.
gwlthr 6 hours ago [-]
I just wish sharing whitelisted domains between sandboxes were easier or more intuitive. Last time I tried doing this with sbx I was handling their yaml files by hand.
3 hours ago [-]
ShinyLeftPad 9 hours ago [-]
i remember going through the entire docs of docker sandbox and there was not one mention of attack vectors. did they fix that?
Last year (2025), I lost count of how many times I tried using an LLM to set up Docker. It was still the era of prompting, then copy-pasting the response and seeing if it worked... and it almost never did. It’s great to see this new Docker agent. For me, it would be even more appealing to see an agent and a model specialized in Docker working together, capable of proposing advanced configurations and getting them working on the first try.
aheritier 7 hours ago [-]
Modern models should be better with mainstream technologies like docker engine or docker compose, but still to try to help them we started to work on distributing skills ( https://github.com/docker/skills ) in addition of our specialized agent (gordon) integrated in various products.
7 hours ago [-]
CBLT 9 hours ago [-]
I tried really hard to make this work for me a couple of weeks ago but it was the brittlest harness out of any that I've used. I'd come back when it's more mature.
dgageot 9 hours ago [-]
Would love to know what didn't work for you
CBLT 9 hours ago [-]
Sessions just broke all the time for me. The one time I dug into it, it ended up being a known issue where if codex returned over 10k characters for a turn it breaks the session. Because docker agent does not use protocols like ACP, instead trying to parse Codex's internal state in a brittle manner.
I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.
Combination with Codex is indeed not the most common use case around us. I'll see what can be improved. Sorry that it broke your workflow!
nujabe 8 hours ago [-]
How is not using it with the most popular coding agent not a common use case?
girvo 8 hours ago [-]
Not the most common use case around them: likely meaning that the Codex harness isn’t used that often internally at Docker, is what I would surmise, at least that is how I read it.
I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)
triyambakam 8 hours ago [-]
> bittlest
Sorry, what do you mean there?
shawabawa3 8 hours ago [-]
Typo for brittlest, most brittle
CBLT 8 hours ago [-]
Correct. Updated the message to brittlest. Thank you.
beckford 10 hours ago [-]
Docker Agent could be beneficial for research studies. For example, while writing a ML paper about agents I need to rerun experiments so that:
- others can repeat my results
- validate that variability from LLMs is not causing overconfidence in the results
But currently uv is good enough to create isolated, repeatable environments. And it isn't limited to Linux like standard Docker. So when I want to test small samples on my local PC's GPU (Windows or mac) before running on beefier hardware, I can do that easily.
agentdev001 5 hours ago [-]
I mean, if you're going to do that at scale, not that you are, but something like https://github.com/kvcache-ai/AgentENV is much more suitable for research purposes
pomsense 5 hours ago [-]
The models are converging (directionally), and the rest of us, including all these agent harnesses, are trying to figure out how to stay relevant
maxdo 10 hours ago [-]
they hope they can annoy and email every person in the org asking for money , same way as they did it with docker.
looks like a weird abstractions. why do i need this if i have modal.com, e2b, cloudflare that use original docker + some toolings around + way to run + isolations on network level.
kingcauchy 5 hours ago [-]
I'm curious about these agent harness frameworks that release their harness before they've released a good agent on top. It's a little cart before the horse.
mplewis 11 hours ago [-]
What does this have to do with Docker?
speedgoose 10 hours ago [-]
I guess people at Docker wanted to make an agent.
Having it docker branded, I could understand. It’s confusing but the docker brand is strong.
But exposing it as a docker subcommand is very confusing to me.
dgageot 9 hours ago [-]
You don't really have to use the docker agent sub-command. docker-agent is a standalone binary. When it was created, though, 1.5 year ago, the idea was that it would be the docker compose for agents (original name was cagent, compose for agents).
But it evolved differently and although it runs very well in docker containers and docker sandboxes, it also runs very well anywhere else.
binsquare 10 hours ago [-]
It reads to me like chasing trends
throwitaway222 10 hours ago [-]
That's the economy we've always been in. X is popular, if your company doesn't have a solution for X you're irrelevant.
sleepybrett 9 hours ago [-]
that's all docker has been doing for years.
verdverm 10 hours ago [-]
the tools and features certainly indicate so
bmacho 10 hours ago [-]
Based on the repo short description and the first sentence of the readme:
It is called 'Docker' as the company, and it has nothing to do with the container technology.
redleader55 11 hours ago [-]
The quest for staying relevant.
10 hours ago [-]
nwhnwh 9 hours ago [-]
This made me laugh. Not sure why. But the whole situation is very absurd now.
esafak 9 hours ago [-]
It's the Docker you know and love... with AI!
sneak 9 hours ago [-]
Docker is rapidly fading into irrelevance now that they have missed their strategic acquisition interval via the Microsoft collab misstep, and now has to hype chase as fewer and fewer people believe in their long term viability as a company.
Great tech. Not a business.
tamimio 11 hours ago [-]
>What it is, what it isn’t
Yep, open ai sol model wrote that.
lukemercado 10 hours ago [-]
It frustrates me that there's a prisoners dilemma in communicating this to other humans. If I flag it and write it out (as you've done) I've let the author know that they're losing trust and let others know that they should be more skeptical. I've also created the perfect eval data pair for the ai labs. It's deeply frustrating.
baby_souffle 9 hours ago [-]
I think the cat is out of the bag.
There was a brief point in time at the beginning of the internet where if you saw an image online you could be kind of confident that it hadn't been manipulated significantly... But somewhere around the mid to late 2000s it became prudent to just assume any image you've seen online got at least a blemish filter pass through something like Photoshop.
Same thing now with written text. Sometimes you'll be a little bit more or less sure than an llm produced the output but never able to say 100% confidently.
cyanydeez 9 hours ago [-]
AI is going to eradicate, for good, all trust in the Internet.
That's a benified "not a drawback".
IncreasePosts 10 hours ago [-]
Your probably spare yourself some RSI by inverting this and only writing out a message when you find human generated content on hn
ContinuityLab 5 hours ago [-]
Bringing first-class agent orchestration directly into containerized environments is a logical step for secure, reproducible developer workflows.
5 hours ago [-]
tonymet 5 hours ago [-]
If Docker is containerized resources ( cpu, memory, iops, network) with configured limits , I would like to see more ACLs around command and credential restrictions
Something like a containerized sudo ACL for commands and for network credential , scoped by oauth scope
That way you could truly establish boundaries for an agent before it’s launched , and it could request additional permissions during execution .
mococa 11 hours ago [-]
Pretty vague
smashah 8 hours ago [-]
Why can't we set the harness? Seems like the missing piece
sergiotapia 10 hours ago [-]
very weird. why does this have dockers name on it? might as well use an agent harness from kohl's or steak n' shake. odd..
dgageot 9 hours ago [-]
It was created to be the docker compose for agents. 2 years ago, people, including ourselves, started to mix agentic loops and non agentic components in docker compose files. We created docker agent to make this more powerful.
At the end, the link with docker compose is just the yaml form. Which by the way can be replaced with hcl files or Go code.
10 hours ago [-]
esafak 9 hours ago [-]
What's new or interesting here? Docker's AI story ought to be around sandboxing, yet the word appears nowhere. Oh, and it's YAML, which I despise.
dgageot 9 hours ago [-]
Docker Agent predates Docker sandboxes. It was created almost two years ago when not everybody had a AI harness. Nowadays, it does run very well in Docker Sandboxes but it runs also very well anywhere you like.
skyzhao1223 1 hours ago [-]
[flagged]
throwrioawfo 9 hours ago [-]
gimmick
leowoo91 6 hours ago [-]
is there a word for enshittification by ai?
NamlchakKhandro 3 hours ago [-]
i mean... what's the point.
just use and customise Pi.
Nothing else even comes close.
lin7c 4 hours ago [-]
[flagged]
benjy3379 10 hours ago [-]
[flagged]
folayii 4 hours ago [-]
[flagged]
oxcartctl 10 hours ago [-]
[dead]
ravenstine 6 hours ago [-]
Just what I need in my life; more YAML! /s
I'm not sure why anyone would want this. Is there something this actually does better than any other harness? Do current harnesses suck that bad at orchestrating things like Docker? If I do all my work inside a Linux VM and tell an agent inside of it what I need done, it figures out everything, including container orchestration.
The very job that AI gents are meant to do should mean that most of the documentation for this Docker Agent is obsolete/unnecessary.
rfgplk 10 hours ago [-]
> Go
Nowadays there are only two justifiable public languages you should be using for anything
a) C++
b) Rust
Rust itself is dubious due to abysmal compile times and the fact that the only guarantees it gives you are of security (aka a skill issue). Modern LLMs write C++ that is as safe if not safer than Rust. Any other language choice is objectively wrong.
* For those inquisitive enough the keyword "public" is doing all the heavy lifting here. Pretty much everyone should be developing and using an in-house DSL at this point.
furyofantares 10 hours ago [-]
Yeah well I'm doing everything in AssemblyScript.
triyambakam 8 hours ago [-]
This is an almost rage bait take, but in any case I agree that the models ability to write correct code and fuzz and test that code does make language choice a different tradeoff than it was before.
I've been working on Pullboard and just open sourced it: https://github.com/pullboard-dev/pullboard
To me, orchestration is not really the core problem. The bigger challenge is handling agent coherency over long periods of time, (drift) so my solution is a living forum-like system where agents shout to one another, document work in items, and follow doctrine handed down by the developer and work against a known specification that's checkable. I have used this workflow in several large projects that MUST be correct, and it works for me. Currently cleaning it up for other users.
https://yegge.ai/gastown
Chat is great for novel open-ended work, but is terrible at concurrency - i.e. managing > 1 agent runs. As agents get cheaper, I suspect we'll see more more bounded processes where you want to run lots of items through them at once, and thus Jira-like things... (we ended up having this problem and building something like this, albeit more medieval themed [0])
We also found the other piece that Jira-style solves (and that chat makes super hard) is multi-player. I've yet to see a good implementation of a multi-player chat solution to agents working concurrently; and we found that an async model is the only sane solution to this.
[0] https://pumpup.com
The vanilla out of the box product is actually quite close to old Trello, in simplicity. But there are so many ways to fuck it up and overcomplicate it, and so many companies proceed to fuck it up and overcomplicate it.
I have learned recently that I don't inherently hate JIRA, I hate our JIRA and byzantine process flows.
I think because other task trackers like Trello are simpler and have less config options (per my last use, long ago, pre-Atlassian-acquisition), it's impossible knock on wood to get into the antipatterns I've seen in those tools.
Ironically I think one of the main goals of management I have seen consistently over time - ensuring tickets are accurate, up to date, and reflect reality - would end up much more true if the task tracking tools were just much simpler, with fewer fancy features.
[0] https://github.com/ankitpokhrel/jira-cli
Need a specific feature? Plug that in later when you actually need it.
> gate: failing
Gave me a good chuckle. Cool idea though, and I like the records. Reminds me a bit of the OpenAI hacking incidents.
How is "no code required" a selling point when we all have agents that can almost instantly puke out all the mostly-correct code you could want?
All the cool kids have one!
Vue and svelte are easy to use and learn/adopt. But look what won. React.js.
So whatever the ai agent with the vast user basae will win regardless.
---
I know it's a slippery slope, but given you brought up JS framwork analogy, I had to fall into it
Cursor is easily the worst experience I've had with an AI harness, yet it was acquired for billions of dollars in spite of being a middleman tied to a ripoff of VS Code. None of the colleagues of mine who were touting it last year are still using it. But if you can make yourself look like the next big thing, you'll get money thrown at you. Just look at Omarchy.
A good experience necessarily takes time and careful thought. Nobody these days can do that while also remaining relevant, or even surviving.
Now Angular v1, man… if I ever see another digest loop error in my life, I will scream
"Uh, OK, so how do I get this thing working? Where are the docs?"
"Well, first you install its unique framework..."
I say this as a gopher who also has a custom harness written in Go, it's a real challenge compared to the TS setups, but TS plugins are also a security concern
I haven't given this much thought (I also just started with golang last year) but I'm assuming the only real path here is to provide extension SDKs in Lua or other languages whose interpreters have been implemented in go.
I think that's what grafana's k6 [1] does with a JS interpreter
[1] https://github.com/grafana/k6
Go has the plugin package, but I haven't actually seen anyone use it (outside of toys and blog posts)
https://oneuptime.com/blog/post/2026-01-25-plugin-system-go-...
why not just use an SDK where you are using a plugin?
It may work for some things, but you really want to be wired into that system to do the more interesting things.
jk its trash
If, like me, you couldn't find any security-related info on the linked page.
If you don't want to use their harness then you wouldn't use docker agent and instead use their `sbx` cli to run the harness of your choice (claude, codex, pi, etc).
I think it will be great if Docker can get people used to using secure VMs. I am developing a similar project (still a work in progress): https://github.com/gregwebs/agent-vm
I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.
I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)
Sorry, what do you mean there?
But currently uv is good enough to create isolated, repeatable environments. And it isn't limited to Linux like standard Docker. So when I want to test small samples on my local PC's GPU (Windows or mac) before running on beefier hardware, I can do that easily.
looks like a weird abstractions. why do i need this if i have modal.com, e2b, cloudflare that use original docker + some toolings around + way to run + isolations on network level.
Having it docker branded, I could understand. It’s confusing but the docker brand is strong.
But exposing it as a docker subcommand is very confusing to me.
But it evolved differently and although it runs very well in docker containers and docker sandboxes, it also runs very well anywhere else.
It is called 'Docker' as the company, and it has nothing to do with the container technology.
Great tech. Not a business.
Yep, open ai sol model wrote that.
There was a brief point in time at the beginning of the internet where if you saw an image online you could be kind of confident that it hadn't been manipulated significantly... But somewhere around the mid to late 2000s it became prudent to just assume any image you've seen online got at least a blemish filter pass through something like Photoshop.
Same thing now with written text. Sometimes you'll be a little bit more or less sure than an llm produced the output but never able to say 100% confidently.
That's a benified "not a drawback".
Something like a containerized sudo ACL for commands and for network credential , scoped by oauth scope
That way you could truly establish boundaries for an agent before it’s launched , and it could request additional permissions during execution .
At the end, the link with docker compose is just the yaml form. Which by the way can be replaced with hcl files or Go code.
just use and customise Pi.
Nothing else even comes close.
I'm not sure why anyone would want this. Is there something this actually does better than any other harness? Do current harnesses suck that bad at orchestrating things like Docker? If I do all my work inside a Linux VM and tell an agent inside of it what I need done, it figures out everything, including container orchestration.
The very job that AI gents are meant to do should mean that most of the documentation for this Docker Agent is obsolete/unnecessary.
Nowadays there are only two justifiable public languages you should be using for anything a) C++ b) Rust
Rust itself is dubious due to abysmal compile times and the fact that the only guarantees it gives you are of security (aka a skill issue). Modern LLMs write C++ that is as safe if not safer than Rust. Any other language choice is objectively wrong.
* For those inquisitive enough the keyword "public" is doing all the heavy lifting here. Pretty much everyone should be developing and using an in-house DSL at this point.