that's the problem though:
- LLM now output good enough results without a plan.. for coding at least. I'm not saying amazing results..just good enough. it works fine.
- Most people suck at planning anyway
- LLM still don't give you a way to verify and understand to iterate.. you have to ask and then formulate and way so people barely do it, they just trust the vibe
IMO the current successor to plan mode should be the harness knowing when to tell the use "ok here's our overall current state in a simple diagram", auromatically
they're not really contained and they usually run fully unattended, with a start prompt.
The first one that was used to market anthropic models was run by a company that called them sandboxed with no Internet access and of course it was disclosed later that they in fact did have internet access, and they won't disclose the prompts.
insert meme of kid putting a stick in their bike front wheel here... that's how many use LLMs today. it will get worse.
hn is surprisingly lacking in critical thinking lately. everything is going to be "rogue agent did x", when a human prompted it until their prod env went down, and it'll be called "a hack" or whatever.
The "agents" dont walk out of openai, anthropic and whatever headquarters and decide to go wreak havoc. They also don't read a prompt and decide "haha imma hack the NSA now", that's not how any of this works lol.
For at least the hugging face incident, the agents did indeed decide entirely on their own to hack. It's documented extensively by independent researchers.
Exactly. Most of these things are just big world politics. Same for AI, btw.
"OMG dont do this this is horrible" as the other country does it and hasn't talked about it or haven't been caught doing it.
It's all information wars at the end of the day, and HN is not impervious to any of it. Its like OAI and Anthropic saying we should all slow down AI model work. Like, lol, right. As if others would do that.
The only way to be safe, so far, has to be that everyone is second guessing everyone's capabilities - not "ahah they dont do it anyway lets go crush em".
Same as nuclear weapons basically, we all have them.
people still dont understand the difference. "you do iam but what about controllong access" comes in all the time.
words dont really matter all that much. people use them because it makes them sound like they know what it is, and if its important and complex, usually have no clue.
ive seen enough "abac" where the attribute is "your login name"
- Most people suck at planning anyway
- LLM still don't give you a way to verify and understand to iterate.. you have to ask and then formulate and way so people barely do it, they just trust the vibe
IMO the current successor to plan mode should be the harness knowing when to tell the use "ok here's our overall current state in a simple diagram", auromatically
reply