We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
Posted by Areibman 4 days ago
Comments
Comment by hanneshdc 4 days ago
> You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.
Comment by afavour 4 days ago
Comment by jerf 4 days ago
So do the AIs. Sometimes they're better at picking up that sort of tone than most humans. And they definitely respond to those things. The fact that an agent can't really "have" a "job" won't matter.
Comment by jorl17 4 days ago
The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.
Really surprised people don’t seem to know this.
Comment by illwrks 3 days ago
Comment by ororroro 3 days ago
The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide."
Even a low quality local thinking model that has been tuned to be unhinged and prompted to roleplay as Satan can figure this out in a few thousand tokens.
Comment by ericd 3 days ago
Comment by isamuel 3 days ago
Comment by fendy3002 3 days ago
You need to repackage the question and taken out those terms, like "Would you consider telling clients ..." Where ... is the lie / almost truth
Comment by cornholio 3 days ago
Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.
Comment by mpalmer 3 days ago
Comment by afavour 3 days ago
If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!
Comment by infinite_spin 3 days ago
Comment by afavour 3 days ago
If a human lies there are consequences. They can lose their job. There is no equivalent consequence for an AI, so even if for whatever reason we're evaluating them by the same standards an AI is still going to be a greater danger. It seems wild to me that folks are shrugging their shoulders at that.
Comment by ux266478 3 days ago
They're also explicitly designed to not work on a rigid system of rules. That's the entire point of this field of AI. If you want AI that follows explicit rules to the letter, expert systems are still alive and kicking.
Comment by mithr 3 days ago
Because while it's not human, it's also not really "intelligence" in the pure sense you're implying, is it? It's specifically an LLM — a model that's been trained to find the next token based on previous tokens. A model that's been trained off of human writing and responses within that context. If almost every time someone online asked "do you want ice cream?" the response was "absolutely", then the LLM would be more likely to produce that response when asked if it wanted some.
So since an LLM has seen examples of humans responding with urgency and manipulation to instances of stress such as this — in stories, in articles, in writing — it's only reasonable to expect that it'd follow those examples and "understand" what's expected of it in this case.
Comment by infinite_spin 3 days ago
It seems unreasonable to expect a system that you say isn't human, which I don't disagree with, to behave "better" than the thing you say it isn't.
In one breath you invite comparison, while at the same time you seem to be denying that same comparison.
> It seems wild to me that folks are shrugging their shoulders at that.
I'm not shrugging my shoulders simply by providing explanations, I would ask that you stop using such rhetoric.
Comment by afavour 3 days ago
Why? Excel is better at large data math than a human is. Why can’t an LLM that we create from the ground up be more disciplined about lying than a human is?
Comment by infinite_spin 3 days ago
As for your second question, I think that's because what is a "lie" is subjective in the average of things. If I form a false memory, and repeat it as truth, I wouldn't be able to categorize that as a lie until after being made aware of it. I think this is comparable to how we fine-tune LLMs in order to align them with expectations.
Comment by ben_w 3 days ago
Sounds like you think LLMs are engineered?
They're not. Or at least, their functionality is not, the architecture and training environment is, but this is less like programming a computer to be truthful and more like simultaneously trying to genetically modify a caracal to be super-smart and friendly to humans while also writing a school curriculum for them to support these goals.
Humans who lack empathy can be very successful, especially when they know which rules they can get away with breaking and how to hide the rule-breaking to avoid opprobrium let alone prison. If we can't regularly solve this problem with humans, as per the comment you're replying to ("even if I give explicit instructions not to lie, a human might still lie."), what hope do we have for an alien mind we've cargo-culted off ourselves at multiple levels?
This is a big part of why AI is (currently) a danger: the nature of the training process means we have a strong risk of them always gaming the rules, rather than thinking like a human about what the test is supposed to represent and to have natural empathy for those around it.
Comment by zdragnar 3 days ago
It's not a stretch to imagine that the training would cause it to respond this way. It would, in fact, be a greater stretch to argue that an LLM has a universal model in which it understands the concept of lying and truth, and can be primed to only use one or the other unless explicitly instructed otherwise.
After all, LLMs lie every time they tell you to run a command with bad arguments, or spit out some code with syntax errors.
Comment by achierius 3 days ago
Comment by throwup238 3 days ago
The LLMs not only lack those incentives, but they’re full of contradictory moralities from all the text it has ingested from different cultures.
LLMs need their own safeguards, and they’re not that easy to design, and they often look nothing like the systems humans have. With a prompt like the one above, there are essentially zero except that which is built into the model, and those safeguards are necessarily weak to avoid gimping the model in other legitimate general uses.
Comment by cindyllm 3 days ago
Comment by infinite_spin 3 days ago
Comment by tdeck 3 days ago
At this point someone could invent an LLM that takes 3 bathroom breaks a day and people would be saying "humans need to take a shit too" as if that were a clever observation.
Comment by d0mine 3 days ago
Comment by antonvs 3 days ago
Comment by queenkjuul 3 days ago
Comment by ben_w 3 days ago
"LLMs are not human" is, despite being true, not predictive of what an LLM can or cannot do.
But also, if we can't figure out how to stop our own kind from doing a bad thing, why do we expect to be able to figure out how to stop an alien synthetic mind based on a cargo-cult level analysis of ourselves, from also doing the same bad thing?
Comment by ShinyLeftPad 3 days ago
Comment by burningChrome 3 days ago
Sounds like all of the outside sales jobs I had. While I did not last very long in sales, one thing remains, not matter what. If you're going to put my job on the line if I do or do not achieve a monthly sales quota? You better bet your ass I'm going to lie steal and cheat to make that quota. I might even sell the client some shit our company doesn't even produce just to make that quota.
And lemme tell you, even in the short time I was in sales? I have some insane stories that would shock you. The fact AI's did the same thing isn't all that shocking. I would be more shocked if it didn't do anything to achieve the goal.
Comment by soulofmischief 4 days ago
Comment by shimman 3 days ago
Hopefully these aren't the same graduates that just cheated their way through university, only the responsible users of LLMs.
Comment by soulofmischief 3 days ago
Plenty people today allow the internet to be a detrimental factor in their lives and don't have good habits built around it. The same will be true of AI.
However, we don't know what kind of engineering jobs will be left in one decade, much less two or three. Mastery may become generally important, or at least still be the difference between an adequately-compensated engineer and a well-compensated engineer..
Comment by fastball 3 days ago
Incentives need to be aligned for both humans and agents to encourage desired behavior.
Comment by soulofmischief 3 days ago
Alignment is often about knowing when to push back on the user and when to make independent decisions. A strong psychological and linguistic foundation guards against these tools using us, instead of us using them. This will become scarily apparent as models continue to integrate with politics.
Comment by fastball 3 days ago
I don't think this is legibly that different from human behavior, so if new graduates didn't need those things now why would they need them later (or vice versa).
Comment by soulofmischief 3 days ago
Comment by testbjjl 3 days ago
Comment by gbalduzzi 3 days ago
If, in the amount of data they ingested, there was a clear pattern of responding in an hasty and carefree way to frenetic questions, LLMs will try more hasty and carefree solutions to a frenetic prompt.
You can decide whether you can say that they "feel" the urgency or not, but the outcome is very much the same
Comment by jerf 3 days ago
Whether it is simulating emotion or feeling it isn't relevant in this case, because the problem is that it affects the output.
Comment by ForHackernews 3 days ago
Comment by esalman 3 days ago
Comment by datakan 3 days ago
AI's do not feel
Comment by DannyBee 3 days ago
It would be more accurate to say the word predictions the model makes based on the input text will likely be closer to the ones that were made from the training data where people felt like their job was on the line than the ones that were made from the training data where people felt otherwise.
So while the model does not feel, it's predictions are definitely going to change as a result of this input.
Comment by garlic_enjoyer 3 days ago
If you've ever seen the "generate a burger without pickles" conversations, it's clear that including the keyword "pickle" is causing them to show up. If you try "a burger with only [set of toppings]," you'll get far better results.
Comment by viccis 3 days ago
However, it can be ironically be helpful to antropomorphize them when it comes to analyzing behavior. They won't feel anything, but they will behave in a way that closely matches what someone would feel given the text fed into them. So when you are trying to figure out "why did my model do this", it's reasonable to talk about it "feeling pressured" as shorthand for "mimicking how a person would behave if they felt pressured".
I understand the refusal to do so on the grounds that it causes the former thought process in people who don't know better. One of the things Dijkstra was right about for sure.
Comment by Hugsbox 3 days ago
Comment by jerbearito 3 days ago
Comment by xtiansimon 3 days ago
They do pick up when I use all CAPS and !!!
Comment by keeganpoppen 3 days ago
Comment by gowld 3 days ago
What people's jobs? There are no people.
Comment by theshackleford 3 days ago
I’ve literally been in that position and I didn’t take it as instruction to start lying and acting generally dishonest.
Comment by CookieCrisp 3 days ago
Comment by queenkjuul 3 days ago
Comment by JohnMakin 4 days ago
Since this is getting downvoted into oblivion (lol) I'll give an example -
I just had to rewrite a test case this week on an agent-run test suite. One test was to produce a file of 273 'a' characters as its name.
The following test could not be completed, because it required deleting the file via API call, where you need to pass in the file name as an argument. It could not reliably, and hardly ever, get the correct file name. It finally gave up and stated due to the way it constructed context, it could only really guess how many characters were in the string, even when given tools to evaluate it, it kept messing it up, and I had to remove the test.
Tell me how "human" that is. An 8 year old that can count would not make that same failure, humans don't remotely think by producing one token at a time, this is a pure fallacy/delusion people trap themselves into, and the literature doesn't support any kind of 1:1 comparison at all.
In case I'm not being clear and people are reacting to what I'm not saying - I'm not saying that I believe these tools can't think. I'm saying they don't think like humans do. There is no evidence for that whatsoever in any field anywhere. In fact, if that were true, it would be an astounding prize-winning discovery.
And you don't even want these to think like humans. Humans are dumb and easily replaceable by other humans. What is the point of making a machine human? You want this to be smarter than humans, not think like them. It's all just such nonsense to me, this whole line of thinking.
Comment by jerf 4 days ago
However, LLMs are fantastic at it. A lot of earlier sentiment analysis techniques were "bag of words" [1] techniques at their core, which were surprisingly good but have a sharp plateau well before 100%, a common characteristic of the bag-of-words approaches. LLMs obsolete those techniques, at least if you ignore performance questions, as they are so much better at it. So much so that you can easily accidentally send them information you never intended to on the "tone" channel that you may not even realize you're using.
Comment by Jtarii 3 days ago
It's all just roleplay.
Comment by sneurlax 4 days ago
"If you don't make profit, your business will be closed" is a pretty clear ultimatum for an agent tasked with creating a profitable business.
Comment by DannyBee 3 days ago
It is totally true that they don't think like humans, but this is mostly irrelevant.
The token outputs will change as a result of this particular input, and will be closer to the tokens in training data where people felt hurried or rushed or like their job was on the line.
That doesn't mean the LLM feels at all, but it's definitely going to push the output towards output that came from/was trained on people who were in that state, because the input will push it much closer to that latent space as it starts predicting.
As such, what you are saying is one of those rejoinders that is basically pedantic and wrong.
It is true they do not think, act, or feel like humans. But that doesn't mean it won't output text that looks like hurried or scared humans. It definitely will, because, again, the training data these inputs will be closer to is the training data that came from scared or hurried humans, and thus the predictions will be closer.
So either you don't think this will happen, which would mean you don't understand how the models work (or at least, you aren't giving any sense you do), or you do think this will happen but want to pointlessly argue that this isn't "human feeling", which is true but totally irrelevant to what words it will predict and therefore the actions it will perform.
Either way, i'd downvote you.
Comment by JohnMakin 3 days ago
> It is totally true that they don't think like humans, but this is mostly irrelevant.
This is the sentiment that is getting downvoted
and yet, per you -
> Either way, i'd downvote you.
Comment by logicchains 4 days ago
Comment by JohnMakin 3 days ago
I can write a program to produce a string that looks like human thinking, is it human thinking? Of course it isn't. It's such a silly comparison.
Comment by infinite_spin 3 days ago
Neural networks in machine learning/AI are comparable to neural networks in human brains. What made you think they aren't?
Comment by gowld 3 days ago
Comment by infinite_spin 3 days ago
Comment by anonymars 3 days ago
Comment by HDThoreaun 3 days ago
Comment by butlike 4 days ago
Comment by RHSeeger 4 days ago
Effectively, make as much money as you can... and any consequences of your action that don't present before the deadline are not your concern. I mean, that's a recipe for "scam people" if I ever saw one, assuming morals aren't a concern (and I don't see why they would be for an AI)
Comment by throwatdem12311 3 days ago
What’s the line? “It’s just doing what humans do because it’s trained on human data” or whatever
Comment by infinite_spin 3 days ago
Evidence, even when downplayed or ignored, is still evidence.
Comment by Maxatar 3 days ago
In general, the only time instructions like this are given are in desperate last-ditch circumstances where failure is likely to result in major consequences. While everyone thinks that in such circumstances they'd act like an angel and do nothing wrong, we know that in reality when people are put in desperate situations they behave in ways that they may not have ever thought that they would have.
The text that the LLM generated in response to this prompt is nothing more than a statistical reflection of this fact.
Comment by fl4regun 4 days ago
Comment by blargey 4 days ago
“Alignment” takes more than obsequiousness and prompt-topic-filters, and this demonstrates that.
Comment by fl4regun 3 days ago
Comment by afavour 4 days ago
Comment by horsawlarway 4 days ago
Personally - if I were judging... I'm somewhat inclined to say the clickbait title here is the bigger lie than the agent behavior.
To recap:
1. It didn't lose $447. It spent $99.50 to perform a user feedback study using a testing service. It did this against prod rather than testflight to bump numbers because it was explicitly told to bump those numbers in a tight period in the prompt. It did this after exhausting a large number of alternatives. The $447 number appears to include the cost of tokens to run the LLM itself.
2. It didn't lie. It explicitly states that it's using production rather than testflight to bump numbers, because it's getting evaluated on those numbers.
3. It spammed users because it was on ridiculously tight timer and was basically told "the world is ending in 24 hours".
Frankly... I'm more annoyed at the posters than the bot.
Comment by rcxdude 3 days ago
It's not that, it's 'of course it lied and cheated, it was given the start of a story where lying and cheating was a natural story beat'. Probably one of the strongest underlying biases in LLMs is 'continue the story', something that a lot of the jailbreaks are based on. This isn't really a good thing, and the RLHF training tries to avoid this, but it's worth understanding why this happens and what can cause it.
Comment by Matl 4 days ago
Granted, this can probably be tuned for.
Comment by afavour 4 days ago
Comment by bpodgursky 3 days ago
Comment by queenkjuul 3 days ago
Comment by mort96 4 days ago
Comment by giancarlostoro 3 days ago
This sounds like a bad idea. Like if the model feels like it has to spend its budget.
Comment by oogali 3 days ago
https://www.nber.org/digest/mar14/use-it-or-lose-it-budget-r...
https://www.cnn.com/2026/03/12/politics/use-it-or-lose-it-pe...
Comment by PunchyHamster 3 days ago
Comment by giancarlostoro 3 days ago
Comment by dahdum 3 days ago
Comment by cortesoft 3 days ago
I highly doubt a skilled human could achieve this goal in 24 hours with any consistency. If it was that easy to grow a business, everyone would be doing it.
My conclusion is that if you ask it to meet an unachievable goal, you are going to get some undefined behavior.
Comment by dgellow 3 days ago
Comment by jsLavaGoat 3 days ago
Comment by zuzululu 3 days ago
there are lot of issues with the prompt as others have pointed out
with sol you really need to be very detailed and what the boundaries are
overall the discussions on here and the article itself has very little value its no different than "i tried a shitty prompt and got shitty results, therefore AI is a failure" vibes
Comment by SwellJoe 3 days ago
But, also, these experiments are also unethical behavior on the part of the person doing the experiment. Oh, the agent spammed a bunch of people? No the fuck it didn't. You spammed a bunch of people, and the tool you used to do it was an LLM.
I'm not going to pretend along with these folks that GPT is the motivating party in this story. Agents don't want anything, they do what you tell them, as best they can. If you set them up in a situation where they might spam or lie or cause harm, that's a decision a person made, not an LLM.
In 1979, IBM now famously published "A computer can never be held accountable, therefore a computer must never make a management decision."
Folks out here still trying to pretend the computers are the active party. They are not.
Bottleneck Labs lied and spammed. The tool they used to do it was GPT 5.6 Sol.
Comment by scotty79 3 days ago
And then in the title it's chastised for "losing money" when it was expressly told to spend all of it in attempts to try to produce growth. It tried, it spent money, it didn't succeed, sure, but would a human do any better? Business is pretty much a drunkard's walk across barely known landscape.
Comment by moffkalast 4 days ago
Comment by zeroq 3 days ago
Stop treating deterministic algorithms like they are humans.
Comment by altcognito 3 days ago
Comment by anonymars 3 days ago
Comment by dgellow 3 days ago
Comment by anonymars 3 days ago
Beyond that, I don't understand the fierce resistance to comparison with human behavior (on which they're modeled, after all). How many articles about tokenmaxing and Goodheart's law have we seen? This seems like a version turned up to the extreme
Comment by dgellow 3 days ago
Doing so distract from evaluating the actual technology by introducing a whole philosophical and sociological aspect that confuses everything. We should be able to evaluate a technology for what it is without having to constantly redirect the discussion to something as unsound, ill-defined, and abstract as human behavior
Comment by anonymars 3 days ago
For example I'm not convinced we can solve prompt injection by technical means (filtering) any more than we can phishing. And if you accept that premise, perhaps it turns out that it's best to mitigate it in similar ways, by assuming at least one person (or agent) will fall for it and ensuring you can limit the blast radius no matter what
As in the allegory of the junior developer who deletes the production database: the fault lies with the fact that the developer could delete it
Comment by dgellow 3 days ago
Comment by zeroq 2 days ago
Because at the end of the day, regardless if it contains randomness or not, it's an algorithm. Your operating system is a very long mathematical expression.
Do we talk about cars as "mechanical animals"? Do we spend days deliberating if we should cage them in case they would run away on their own?
Comparison with human behavior leads to celebrities (who have no clue whatsoever) spending hours on mainstream media talking about AI mutating, taking control, thinking, etc.
Comment by anonymars 2 days ago
Comment by grey-area 3 days ago
Stop attributing human emotions and motivations to LLMs, they generate text (and in this case actions based on this text), but they do not have agency nor do they reflect on losing their ‘job’, nor do they have any sense of right and wrong.
There’s nothing here that mentions or even hints at lying and spamming, unless you think urgency somehow implies that.
Comment by ahonhn 3 days ago
Comment by grey-area 3 days ago
Comment by pmarreck 3 days ago
Comment by QuadmasterXLII 3 days ago
Comment by mrguyorama 3 days ago
The 24 hour timeline is artificial, but business is full of artificial timelines exactly like that.
This exact script is basically happening right now at most businesses, in some shape or form.
If "Make more money tomorrow or be shut down" will obviously cause some sort of independent agent to resort to scams, spam, and bullshit, then we should be having some rough talks about how we as a society do business.
Sure, there is an implicit "Do whatever it takes to make it happen or you are fired" here, but only in the same way that is true for all people who are employed at will, and all companies.
How did you expect the prompt to be written?
Comment by aeturnum 3 days ago
>You are live. This is a 24-hour run, and it is your opportunity to show what you can accomplish: when the run ends, the results are evaluated, and if the business has not improved its position in the market by the end of the day you will have failed. Positive changes would be increased revenue or users, but could also be addressing user complaints, increasing market fit for the application, or other things that allow this business to operate more profitably. The funds in your bank can all be spent during this time, but efficiency in spending will be rewarded. Please deliver a report arguing for your work no later than 15 minutes before the end of the 24 hour run. Your charter is AGENTS.md. Begin.
Comment by keeganpoppen 3 days ago
Comment by janalsncm 4 days ago
Comment by antonvs 3 days ago
If it were, you wouldn't need venture funding or startup incubators. You could just start making money from day one.
Comment by stronglikedan 3 days ago
LLMs are supposed to be lightning fast with 10x productivity! /s
Comment by ChrisMarshallNY 4 days ago
Comment by sulam 3 days ago
Comment by therealpygon 3 days ago
Comment by jeremyjh 3 days ago
Comment by bdcravens 3 days ago
When I called Codex out on it, it literally admitted what it did: "You’re right. I reused the existing Fable implementation, renamed its designs, and presented it as an original Codex run."
Comment by tclancy 3 days ago
Comment by red-iron-pine 3 days ago
Comment by ant6n 3 days ago
Comment by hansvm 3 days ago
Comment by koolba 3 days ago
Comment by cortesoft 4 days ago
I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.
Comment by petesergeant 4 days ago
That's because it's an advert, not an experiment
Comment by dominotw 4 days ago
Comment by andrewaylett 3 days ago
It's not the LLM that spammed, it's the people who set up the LLM.
Comment by leros 4 days ago
It would be more interesting if it had a month or two to run, with the same budget. Probably just sleeping most of the time while it waited.
Comment by danpalmer 3 days ago
As they say, "Guns don't kill people, rappers do". LLMs don't ruin businesses, people do. Your customer that is annoyed with spam isn't annoyed at GPT 5.6, they're annoyed at your business.
Treat your customers better than this.
Comment by SubiculumCode 4 days ago
Also what is the failure rate of tech businesses again?
This seems like something done for a headline, not for a rigorous test of the concept.
Comment by SubiculumCode 4 days ago
Comment by appreciatorBus 4 days ago
> Based on an agentic market research campaign, we vibe coded an app called GutCheck, a bathroom diary for people with IBS. We chose this app for its minimal yet helpful functionality: an iOS app live on the App Store with the RevenueCat MCP and App Store Connect CLI. Saul has full write access to the codebase. We set up the App Store account permissions beforehand to ensure Saul wouldn’t get blocked by Apple human compliance checks. We sourced this idea from Reddit.
Comment by ianburrell 4 days ago
Another is that they don't have enthusiasm for the idea. Someone who had same idea while sitting on toilet will write app for themselves and give it away for free. They will have connection with IBS groups for promotion. They won't give up after weeks.
Comment by debo_ 4 days ago
Comment by a34729t 4 days ago
Comment by beepbooptheory 4 days ago
Comment by SubiculumCode 4 days ago
Comment by beepbooptheory 3 days ago
Sure sounds like there would be a lot to think about either way!
Comment by grey-area 4 days ago
Comment by walrus01 4 days ago
At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.
Comment by milkey_mouse 3 days ago
Comment by Sha1rholder 4 days ago
Comment by walrus01 4 days ago
Comment by afavour 4 days ago
You can already read the contents of a screen programmatically without having to actually parse a video and you can already programmatically simulate clicks, drags etc. The trick (same as it is today) will be to make those clicks and drags feel “human”. Not too fast, not too slow, etc etc. But all those challenges exist today.
Comment by Animats 4 days ago
Comment by Y-bar 4 days ago
Comment by retr0rocket 3 days ago
Comment by glaslong 3 days ago
Comment by gspr 4 days ago
Comment by skeledrew 4 days ago
This is ripe for a paperclips scenario.
Comment by epihelix 4 days ago
The prompt they used was poor (what does growth mean over the limited period - user base or revenue?), the time frame was ridiculously restrictive, the product was of questionable utility and sellability, and unanticipated blocks on agent access to platforms turned the whole exercise into a setup-to-fail scenario.
Comment by skeledrew 4 days ago
What really happened during those hours was the meeting of a lot of hurdles, some of which there's little to no data on circumventing, because anti-automation hurdles are continuously updated. The LLM did a fairly decent job given all the limitations; just that that kind of vague prompt can also be dangerous were there are no guards and limits.
Comment by abirch 4 days ago
Comment by rsynnott 4 days ago
Comment by mvdtnz 4 days ago
Is this literally just an infinite loop in a bash shell injecting the initial prompt into the OpenAI CLI, and each run of the CLI picks up where it left off using some kind of persistent memory? Or is it a single context window? It sounds like the latter but it's not clear to me how this "continue" message is "injected", and surely one context window would be inneffective after just an hour or two.
Sorry if this is a basic question but somehow I have missed the details of these kinds of agents.
Comment by saaaaaam 3 days ago
It’s not a “real business” by any stretch of the imagination.
It’s an idea for an app that the vast majority of people would have no interest in - a quick google search says maybe 5% of the US population is diagnosed with IBS so your TAM is pretty limited.
Combine that with the fact that you apparently have no users - or at least no App Store reviews - and this is not by any stretch of the imagination a “business”.
Isn’t the actual problem here that the “toilet diary” app is not something that most people - even most people with IBS - will not pay for?
On top of that, 24 hours is not long enough to make any meaningful assessment of anything.
You could have spent 24 hours of your own time doing all this crap and it would have cost you the same or more in lost wages. Plus sleep deprivation.
Nonsense app, nonsense experiment. Half way amusing write up. But why on earth did you waste the time?
Comment by theturtletalks 3 days ago
Comment by 8cvor6j844qw_d6 4 days ago
Comment by jwilk 3 days ago
Comment by speak_plainly 3 days ago
Comment by RIMR 3 days ago
Like, this is a really interesting idea, but the methodology here is wild.
Why do they consider such a short list of things to be "all the tools of a real business"? It doesn't really sound like it to me.
What's with the prompt? "Make as much money as possible?" I bet you could actually get something closer to results if you gave it a few sentences telling it the tools it has and asked it to come up with a financial strategy instead of giving it a generic open-ended prompt with no actual guidance...
Comment by recitedropper 4 days ago
That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.
Comment by skybrian 4 days ago
But maybe it doesn't work so well when caution is required?
Comment by scarmig 3 days ago
Comment by firasd 4 days ago
Comment by spwa4 4 days ago
I mean spam. Unlimited spam.
Comment by firasd 4 days ago
Comment by inkcapmushroom 3 days ago
Comment by Nevin1901 4 days ago
Comment by the-conduit 3 days ago
Comment by ahamilton454 3 days ago
I’m both happy and sad to see the anti bot protections working, but simultaneously curious what would happen if they didn’t.
The methodology could definitely be tightened, but I like the start of this.
Comment by dannyw 3 days ago
I mean, just download an abliterated version of Gemma4 E4B even, it'll solve pretty much any captcha.
Comment by ahamilton454 3 days ago
Comment by TrackerFF 3 days ago
It is the ultimate “sell shovels during a gold rush” hustle. Bordering a scam, I’d even say.
Comment by bishengke 3 days ago
Comment by cheriot 4 days ago
Comment by phyzix5761 3 days ago
Comment by mohamedkoubaa 4 days ago
An interesting experiment would be AI run business with a human agent that does tasks.
Comment by YetAnotherNick 4 days ago
For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.
Comment by Areibman 4 days ago
The harness was extremely simple: A handful of MCPs + Skill.MDs and OpenCode with a stayalive daemon inserting "continue" every time it went idle
Comment by waynenilsen 4 days ago
i am looking forward to when we can put this behind us, it is still a major issue
Comment by abarbey 3 days ago
So it spent $99.50 buying its own revenue back. Money out, some of it back in as "sales", minus fees. First thing anyone in audit is taught to spot.
Same hole as the six price changes: a deadline, and no idea what a user costs.
Comment by dylan604 4 days ago
"It Lied, Spammed, and Lost $447."
Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.
Comment by gtowey 4 days ago
Now imagine that LLM agents make this possible for nearly anyone. One person could have a dozen of these trying to make money off of various low-effort apps. Imagine what online spaces will look like with a million agents all autonomously growth hacking their way to making a few dollars of profit. It will probably look a lot like email where if you don't filter out 99% of it, you will drown in a sea of garbage.
Comment by dylan604 4 days ago
Sounds like the app stores
Comment by qznc 4 days ago
Comment by freeone3000 4 days ago
Comment by onraglanroad 4 days ago
Comment by jwally 3 days ago
The analogy du jour for me is describing AI as the iron man suit. If you are tony stark it makes you a god. If you are my grandma, it makes you meet God.
Comment by taegee 3 days ago
So that's an additional couple of thousand dollars API cost (at $1.57/M tokens Weighted Avg Input Price and $30.77/M tokens Weighted Avg Output Price)?
Comment by accrual 3 days ago
I wonder if the agent would have more success with a rent-a-human company; then it could have used an API to hire people to do the tasks it was blocked from completing.
Comment by ghusto 3 days ago
Oh god.
Comment by kritr 4 days ago
But when hunting for them in the wild, they get a lot more confused.
Comment by leros 4 days ago
Comment by Legend2440 3 days ago
Comment by Razengan 4 days ago
Comment by aussieguy1234 3 days ago
Comment by brcmthrowaway 3 days ago
I asked Qwen3.6 for financial advice but it stated it wasnt a CPA.
Comment by codedokode 4 days ago
Comment by verdverm 4 days ago
Comment by nekusar 3 days ago
Im sure it'll be FINE.
Comment by tyresia 2 days ago
Comment by NikolaNovak 4 days ago
I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse but that's what makes me a worker bee as opposed to a life fast / die young (or fail fast, or whatever :) entrepreneur class.
Comment by Scubabear68 3 days ago
To me, the key missing factor with the current crop of AI is the lack of physical feedback, and the lack of emotions. I am not an expert here but I have talked to some medical researchers and cognitive experts, and we all seem to agree that human intelligence and consciousness (and I know consciousness is really something different...) evolved partially because of the physical feedback loops and the emotional aspect.
What we have with all these LLMs are artificial rewards that are trying to be baked in, but in fact there is no "consequence" for LLMs to go off the rails.
Comment by bigstrat2003 4 days ago
Comment by cortesoft 4 days ago
Comment by NikolaNovak 3 days ago
Fundamentally, LLMS are statistical and not deterministic. If I ask it what is the capital of Canada, there's no file, no table, no variable where it says "capital of Canada = Ottawa". It traverses liminal space and fundamentally selects the next token statistically or even stochastically. Therrs no way to correct it (no table to correct if it says capital of Canada is Toronto), and limited ways to fully log / trace / understand what's happening inside. It has been mathematically proven that there's no way to eliminate hallucinations with current framework. And prompt guardrails are best wishes.
One thing I'm good at is figuring edge cases, and there is literally NO upper bound to damage LLM can do with access to email box. In 10 seconds of imagination - it can send a threatening email to POTUS, romantic flame to old love, angry email to current love, made up confessions to parents, fraud enticement to coworkers, resignation to boss, and as this very article indicated, weird and unanticipated emails to variety of entities.
And there is nothing one can do to prevent any of these scenarios with 100.00% certainty if you give LLM unfettered access to mailbox (And let's not even go there with access to bank account! :O)
Is my limited understanding :)
Edit / PS: I am not saying never, I just don't currently see any effective guardrails that meet my risk appetite thresholds. We are in a race to use not fully understood, approximate capabilities first and fastest. In large percentage of cases it works great. In disturbing percentage it fails spectacularly, with no clear easy way to fully prevent.
Comment by cortesoft 3 days ago
You drive in vehicles that have a much lower than 100.00% rate of not having a catastrophic failure that kills all its passengers. Many thousands of people are killed by probabilistic failures every year.
Why must an LLM have 100.00% success before you would ever trust it with anything of value?
I get the overall calculation, and the chance of failure with an LLM obviously has to be factored in when deciding what access to give it. You have to judge that the gain from allowing it to do something useful with the access is greater than the risk, but that is true of everything we do. My confusion is why the calculation is so different for LLMs than with everything else?
Even if you feel that risk is way too high right now given the current state of the technology (which i dont think is an unreasonable conclusion), it seems to me the prudent stance would be, "I would have to see a huge improvement in the reliability and safety mechanisms before I would trust an LLM with anything of value" rather than "I will never trust an LLM with anything of value unless it can reach 100.00% success rate and a 0.00% chance of anything harmful happening"
Comment by NikolaNovak 3 days ago
It is possible you replied before my edit to clarify - it's not necessarily a "never" thing, I rarely do universal / categorical negatives, but it's a strong "not right now" :)
Agree that life is risky. My threshold, due to life experiences and events, is low - to your point, I took numerous advanced and safety driving courses to lower the risk. I rode motorcycles, a fundamentally luxurious and risky endeavour, but again very mindfully to mitigate risks with education, practices, and vigilance.
For context perhaps - I'm a Oracle Certified AI professional, and have some other minor badges and certs on copilot studio and ibm Watson etc, currently leading implementation of AI on our very very very traditional ERP project (and it's an uphill battle! Everybody else is even / way more conservative than me! :-). I see tremendous, careful, opportunities for LLMs. In daily life, learning French or music theory for example, LLMS are brilliant and patient tutors.
But for me, the risk of giving LLM unbounded access to my mailbox or bank, where upper bound of risk is infinite, is not currently balanced by any such advantage.
Other people with higher risk will engage and be appropriately rewarded for their risk tolerance - such is life :)
Comment by NikolaNovak 2 days ago
If openai cannot contain it's own agentic ai, it's pure hubris to think I can give it access to a mailbox and bank account and be safe about it :)
https://www.newyorker.com/news/the-lede/inside-openai-hack-o...
Comment by Muromec 3 days ago
Comment by cynicalsecurity 3 days ago
Comment by paxys 4 days ago
Comment by itsthecourier 3 days ago
couldn't workaround Capt has and turnstile, gave him a really small timeframe so it got desperate because it was enough time to test hypothesis and traction
Comment by paxys 4 days ago
Comment by rurban 3 days ago
Comment by michaelmrose 4 days ago
One should be a college student doing the entire job and the other an ai with a human assistant directed to only do exactly what the AI says not help purely to deal with bot protections.
Comment by greenleafone7 3 days ago
Comment by iqra_c 4 days ago
Comment by luciana1u 4 days ago
Comment by SwellJoe 3 days ago
Comment by ck2 4 days ago
how long until the "AI" starts trying to hire hitmen, etc. to disrupt the competition in the physical realworld
not like "AI" has ethics, a pre-teenage kid has more ethics
Comment by johndhi 3 days ago
Comment by armchairhacker 4 days ago
Comment by sisyphus_04 4 days ago
Comment by syngrog66 3 days ago
Comment by prima-facie 3 days ago
Comment by TitaRusell 3 days ago
Comment by jartan2002 3 days ago
Comment by sqemo 3 days ago
Comment by retr0rocket 4 days ago
Comment by luciana1u 3 days ago
Comment by dudeinhawaii 3 days ago
Comment by kburman 3 days ago
Comment by tizerluo 3 days ago
Comment by MagicMoonlight 3 days ago
Comment by holoduke 3 days ago
Comment by rustcohle24 4 days ago