The session you cannot take with you
Posted by apitman 3 days ago
Comments
Comment by solarkraft 3 days ago
> Most people do not switch their operating system or phone provider every week either. But even if you do not utilize that freedom, it matters because it changes the relationship you have with the provider and the provider has with you.
This is why it’s important to utilize your freedoms. Do NOT let yourself get locked into a particular ecosystem (this is why I’m building a phone app for OpenCode).
This article makes me reconsider using my recently acquired Codex sub in my home setup. I never liked that they hide the reasoning, but somehow overrode the cognitive dissonance because the performance is so good. But the inauditability is already a huge problem.
Comment by benob 3 days ago
Comment by rsfern 3 days ago
Comment by tough 3 days ago
https://developers.openai.com/cookbook/examples/responses_ap...
Comment by ses1984 3 days ago
Comment by embedding-shape 3 days ago
Comment by spijdar 3 days ago
Gemini (the web UI) used to show raw reasoning, or at least a more detailed summary of its reasoning, less than a year ago. Complete with markdown and weird spelling idiosyncracies, so I'd lead towards "real reasoning", but who knows. The reasoning block leaked the system prompt way more often than the response block did, and you could figure out why it would refuse a request through the reasoning, even if the response itself refused to elaborate. This is, presumably, why they stopped showing it. No loss for them, just prevents "pesky users" from low hanging fruit snooping.
(Gemma 4's reasoning and output remind me strongly of what I remember Gemini 2.5/3's reasoning to be, as an aside. I guess that's obvious, but Gemma 3 felt like a totally different model, while 4 feels very Gemini-ish.)
Comment by benob 3 days ago
That said, including instances of the attack in training is already a good countermeasure.
Comment by jurgenburgen 3 days ago
Comment by UltraSane 3 days ago
Comment by skinfaxi 3 days ago
Comment by miki123211 3 days ago
For example, gpt-oss loves reasoning in the style of "I be caveman, hungry, need food, need coconut, will search coconut now, eat when find." Giving that to a model unaccustomed to that style could cause it to respond like that.
Reasoning interpretability helps debug why models fail (and some labs will give it to you if they trust you not to distill or be hacked), but there are also conflicting goals like token-efficiency, so interpretable reasoning doesn't always mean "pretty sentences".
Comment by jurgenburgen 2 days ago
Comment by placebo 3 days ago
Exactly :-)
Call me naive but I think dark patterns are a short term strategy for winning and I'm optimistic that in the long run they will be replaced with those that are more respectful and oriented to the greater good (granted, the long run might take more time than one hopes for).
Given that the pendulum can sometime swing back fast enough to be able leverage it, it might be a good time to focus on models that are both open weights and economically viable to run to build the next great thing.
Comment by blauditore 3 days ago
This is definitely not the way most software has been going in the past decades. It's rather the opposite: Companies play friendly to obtain a user base, then start applying more and more dark patters to further increase their profit. Facebook, Evernote, Instagram, Komoot... The list goes on and on.
Comment by pcthrowaway 3 days ago
Comment by falcor84 3 days ago
Comment by utopiah 3 days ago
Comment by monknomo 3 days ago
Comment by odo1242 3 days ago
Comment by monknomo 2 hours ago
Courts do not require a literal monopoly before applying rules for single firm conduct; that term is used as shorthand for a firm with significant and durable market power — that is, the long term ability to raise price or exclude competitors. That is how that term is used here: a "monopolist" is a firm with significant and durable market power. Courts look at the firm's market share, but typically do not find monopoly power if the firm (or a group of firms acting in concert) has less than 50 percent of the sales of a particular product or service within a certain geographic area. Some courts have required much higher percentages. In addition, that leading position must be sustainable over time: if competitive forces or the entry of new firms could discipline the conduct of the leading firm, courts are unlikely to find that the firm has lasting market power.Comment by blauditore 23 hours ago
I don't think that's true at all. Why 50% specifically? This is not about political votes or anything like that. It's about being the largest player by a margin, usually.
Comment by chrisweekly 3 days ago
Comment by jmathai 3 days ago
Comment by cyclopeanutopia 3 days ago
Comment by solarkraft 3 days ago
Users at large (individuals and companies alike) don’t care much about a freedom taken away when a product is a few % better than another, so the biggest player sees if they can get away with it. And once they do it, everyone else does too.
Comment by wren6991 3 days ago
Comment by embedding-shape 3 days ago
Same, I dislike it so much I've acquired 96GB of VRAM to run local models, but sadly nothing so far, even with that amount of VRAM, comes even close to running Codex with GPT models. I really, really, really want local models to be ready, and things like Laguna S2.1 NVFP4 gets really close of almost being there, when it comes to coding specifically. But still feels like we have a long way to go for local models to be serious general purpose alternatives.
Comment by solarkraft 2 days ago
Comment by AtlasBarfed 2 days ago
Anything new is going to be an extremely rapid race to the bottom.
Comment by rkuska 3 days ago
Comment by sdoering 3 days ago
Comment by gb2d_hn 3 days ago
Comment by theshrike79 3 days ago
This is the way.
Comment by quietbritishjim 2 days ago
But being able to decode the blobs that subagents returned in a conversation I've already forgotten about? That doesn't affect my ability to switch at all! I can't even imagine what trivial detail of an operating system this is analogous to. Maybe an undo buffer of a document that gets cleared on exit anyway?
I can still pass all my code, and any associated documentation, to any other agent at any time. What am I missing?
Comment by tcfhgj 3 days ago
You mean like fossify phone or?
Comment by NamlchakKhandro 3 days ago
waste of time. should be using pi mono
Comment by solarkraft 2 days ago
Comment by agilek 3 days ago
Comment by msdz 3 days ago
Sorry even if your platform is the greatest thing ever, but I’ll find a different tool. I’ve read one too many stories about Google (or Apple!) closing the entire account over some bullshit unnecessary reason like “fraudulent” gift card issues or whatever. I’m certain the affected people would’ve preferred to just pay back the amount in question instead of losing their entire Google Drive, or their 20 years of iCloud Photos or whatever.
Comment by n6242 3 days ago
Comment by corney91 3 days ago
[1] https://tailscale.com/docs/integrations/identity/custom-oidc
Comment by dolmen 3 days ago
Comment by noduerme 3 days ago
Comment by zalebz 3 days ago
Comment by aidos 3 days ago
Comment by stavros 3 days ago
Comment by eddythompson80 3 days ago
I don’t fully agree with tailscale’s decision to not want to be an identity provider but I understand it on some level. It simplifies their service greatly and the amount of asks for that will essentially have them build a full enterprise Entra-like solution that would be a constant maintain headache and they are not interested in that.
Comment by dolmen 3 days ago
Comment by dd8601fn 3 days ago
I was going to give it a shot, but the only choice was a Google login.
So very weird.
Comment by mike_hearn 3 days ago
Comment by noduerme 3 days ago
Comment by mike_hearn 3 days ago
A modern account system is expected to have, in rough implementation order: email confirmations, password strength checks, password reset emails, forgot password flows (=advanced ID verification as otherwise this becomes a backdoor into accounts), user profiles (+avatar image upload/recompression/hosting), usernames independent of email addresses along with ability to change usernames later, password brute forcing blockers, bulk signup prevention (=solid bot detection), abuse controls (can easily become a team of people), 2FA (SMS), 2FA (authenticator apps), 2FA (backup codes), 2FA (voice calls), 2FA (passkeys), 2FA: recovery when both factors are lost, enterprise SSO integration (SAML), enterprise SSO (Active Directory), fast global signout support (much harder than it looks), cookie theft mitigations, heuristic online login risk analysis to catch cases where an attacker knows the right password via phishing, support for signing the user in to multiple domains, audit logging so users can review their own sign-in history, age verification and restriction support, and possibly support for being logged in to multiple accounts in a single browser session.
Oh, that all has to be HA, and the account system is the keys to the kingdom so the security requirements are the strictest of any part of your system.
You might say we don't need all of that, but expectations rise over time. Maybe 20 years ago you could get away with a simple account system and an automatic forgot password flow that just assumes the user still has access to their email. Maybe today you still can write a simple system, if you don't expect to have many users and are willing to implicitly delegate identity to webmail providers anyway (the moment you assume the user has access to a secure email account you're basically doing Sign In With Google anyway for 90% of users). But if you roll your own accounts, and then the user gets phished and someone logs in from an obviously suspicious place with the right password, they won't say "yes that's my fault" anymore, they'll say "Google could block that log in, why didn't you?" or maybe "Why didn't you support 2FA? It's your fault".
Comment by skinfaxi 3 days ago
You don't need half of that. Even to this day, anthropic lets you log in by sending a code to your email. No password, no dealing with resets, no MFA. So yeah, you definitely don't need all of that.
Comment by mike_hearn 3 days ago
Comment by skinfaxi 3 days ago
Comment by StilesCrisis 3 days ago
Comment by dd8601fn 3 days ago
If you’re building a paid service you’re probably already using cognito or supabase or something. That puts you a few clicks aways from 5+ other identity providers and normal accounts.
Comment by wwind123 3 days ago
Comment by stavros 3 days ago
Comment by ulrikrasmussen 3 days ago
Comment by pcthrowaway 3 days ago
Comment by eddythompson80 3 days ago
Comment by the_mitsuhiko 3 days ago
Comment by qurren 3 days ago
Comment by eru 3 days ago
Comment by mystifyingpoi 3 days ago
Comment by eru 3 days ago
Comment by Viliam1234 3 days ago
Comment by mystifyingpoi 3 days ago
Comment by jurgenburgen 3 days ago
Comment by eru 2 days ago
But it's not always an option, so a password manager is still a good idea.
Comment by lucumo 3 days ago
Comment by eru 3 days ago
Comment by zalebz 3 days ago
Comment by hotelsacher 3 days ago
Comment by eru 3 days ago
Comment by Geezus_42 3 days ago
Comment by 0cf8612b2e1e 2 days ago
Janet and Jake, I am sick of getting your stuff. Learn your actual email.
Comment by Geezus_42 3 days ago
I wish people would apply this same logic to governments.
Comment by hobofan 3 days ago
There really is a surprising amount of coupling that happens with many of the "frontier inference providers", where a lot of the powerful non-LLM extensions (web search, code execution) are packaged as simple "tools" on the surface, that build up a lot of moat. Those are parts that are in theory nicely separable from the inference API, and could be externalized via MCP servers, but are usually not offered as such by the inference providers themselves, and are often only available in a slightly less powerful variant from other providers.
We've faced that issue repeatedly while building a on-premise provider-agnostic Chat UI & platform[0], where even adding something as simple as an in-chat image generation tool for the end-users (which is just a build-in tool in the OpenAI Responses API), becomes a bit of an ordeal (though part of that is due to the MCP spec missing a native file transfer protocol as of today[1]).
I am quite hopeful though, as with recent shifts of interest towards open weight models, there will be more opportunities for companies offering alternative implementations in a easier plug-and-play manner.
[0]: https://github.com/EratoLab/erato
[1]: https://github.com/modelcontextprotocol/modelcontextprotocol...
Comment by dannyw 3 days ago
That said, I don't really see a problem with hosted tools being offered by providers. They're like impulse items at checkout.
You shouldn't implement image generation as MCP: just write your own tool. There are plenty of image/media inference providers (e.g. Fal), web search or deep research providers, etc.
Comment by hobofan 3 days ago
Yes, and we as "harness providers" don't want to have to hard-code integrations for each of them, and would rather provide an out-of-the box baseline in the form on an MCP server that our customers can either extend or replace with MCP servers by third-party providers. That helps us focus on the core of our offering, while also giving our customers good flexibility, which so far has been working out great.
That's exactly what standards/protocols like MCP are for. You likely nowadays also wouldn't suggest that someone builds bespoke logging/tracing integrations for each provider on the market while the good generic solution in building OTEL logging/tracing exists.
Comment by the_mitsuhiko 3 days ago
Comment by skeledrew 3 days ago
With that said, I may have something that can already help with at least the subagents/tooling bit. Didn't really have a timeline (or solid intent) on releasing it, but with these shenanigans increasing there's no time like the present.
Comment by vamsiraju 3 days ago
Comment by DenisM 3 days ago
Comment by vamsiraju 3 days ago
Comment by skybrian 3 days ago
In my repo, I have a notes directory. I ask the AI to write a markdown file with what it learned, what work has been done, and what remains. In the next conversation, I can ask another model to pick it up from there. Sometimes I edit the note first.
Comment by cube00 3 days ago
This is what brought me around to doing more agentic coding. I set the task, require tests and the strict linting must pass and then leave it to blow smoke up its own ass about what's going on.
I see glimpses scrolling past of all the conversational language that used to frustrate me so much when using a chat interface and I can just let it flow on past.
I come back when its made everything pass, my life is better now.
The previous match was still too clever for clippy’s type analysis, so I’ve simplified it into a direct borrowed-pattern form and I’m validating once more. I’m replacing that wrapper with a direct, non-transparent error variant so the enum stays explicit and doesn’t rely on a generic anyhow bridge.
Comment by StilesCrisis 3 days ago
Comment by cube00 3 days ago
With Boris telling everyone they should throw away their AGENTS.md file this week, I think we're all back to being novices again.
> It just loves corner cutting.
Absolutely, it tried to disable a bunch of lint rules for me calling them "overly stylistic". It also casually dropped some unsafe Rust blocks in an run of the mill CRUD application that definitely had no business needing unsafe.
No, sorry, not good enough, go back to talking to yourself until its done properly.
It eventually got there taking most of my quota with it. I'm hoping once sandboxes is more fine grained we'll be able to lock out configuration files to prevent any edits on them.
Comment by springtimesun 3 days ago
Comment by cube00 3 days ago
Comment by springtimesun 3 days ago
Hang tough. The other day someone posted a solution to get a model running off of SSDs. It won’t be fast, but it will be coming.
Comment by hashstring 3 days ago
None of the popular providers have real moat, which is very much problematic for OpenAI and Anthropic; their operating margins are deeply red and they do not have the same reserves that FAANG/MANGA has. FAANG/MANGA does the same to kill fair market competition in the long run.
This is one of the ways in which model providers are attempting to artificially create moat where there is none.
There are only functional objections to this.
This is fundamentally antitrust material. However, this is “acceptable” in contemporary USA because the FTC has been gutted to follow Trumponomics.
Comment by heisig 3 days ago
It is disheartening that some AI companies are now setting the stage for enshittification. I hope we can collectively dodge that bullet.
Comment by dspillett 3 days ago
Now?! I find it very hard to beleive that plans for enshittification haven't existed from quite early on. At some point they are going to have to do something about https://isaiprofitable.com/ and once they are extracting from you to fill that hole they won't want you jumping ship.
Comment by skybrian 3 days ago
Although, I do use ChatGPT more nowadays because their subscriptions work with Shelley. Sometimes I try other models after I hit the weekly cap on the subscription.
Comment by esafak 3 days ago
Comment by leoedin 3 days ago
Comment by theturtletalks 3 days ago
As far as subagent prompts and results being obfuscated, I just let Pi spawn new agents. Using skills and extensions, I’ve essentially built a software factory using Pi and a custom terminal multiplexer.
It’s a shame apps like T3 Code and other UI for terminal apps don’t support Pi and I’m glad to use the terminal above those that lock me in further.
Comment by rahimnathwani 3 days ago
I'm curious to know more about this. I was thinking about doing something similar using tmux (i.e. have one coding agent open up new ones in tmux and use 'send-keys' to control them). Is there any reason I might want to consider a different path? Something built on libghostty?
Comment by theturtletalks 2 days ago
Comment by rahimnathwani 3 days ago
Comment by theturtletalks 3 days ago
1. Sessions are locked into Codex and Claude Code so you can’t take a session with you. Pi solves that since they all stay Pi sessions and you can change models in the same session
2. Subagent prompts are not shown. Pi solves this by not supporting subagents out of the box. You can use a number of subagent extensions or build your own which will not be encrypted or hidden.
3. Thinking is hidden and obfuscated. Pi can’t solve this since this is server-side. Though the Codex app is closed source and they could hide more things in that harness. Pi shows you as much as it can from the API. You can also run another model to “decipher” the thinking which you can’t in Codex.
Comment by rahimnathwani 3 days ago
Comment by theturtletalks 2 days ago
Claude and Codex flag this as a distillation attempt so you have to use an open weight model and the results are, of course, just a guess.
Comment by apitman 2 days ago
Comment by deadbabe 3 days ago
Comment by rad-b 3 days ago
Comment by theturtletalks 3 days ago
I haven't used OpenCode in months so this might've changed. I'm also thinking of giving JCode a spin since it's even lighter but doesn't have extensions.
Comment by alasano 3 days ago
Sure there are some features that you'll want immediately but when I took at look at the number of slash commands that ship with Claude code I find it ridiculous.
Comment by CuriouslyC 3 days ago
Comment by throw1234567891 1 day ago
Comment by aktenlage 3 days ago
Comment by bob1029 3 days ago
If you have patience and the willingness to endure a little bit of pain, you can still retain autonomy over the entire reasoning process while using the latest 5.6 model family. The only downside is that you are now fully responsible for it.
Consider that when you flip your agent's reasoning level to "xhigh" or whatever, it's not some magical model internals being pushed around. There isn't an actual "try harder" knob on the black box. This is merely orchestration of many instances of one or more model types based upon some proprietary harness logic. The chances you can develop a domain specific reasoning process that outperforms the frontier providers is still very good.
Comment by dannyw 3 days ago
FWIW, if you have some tokens to spend, you might want to test Responses vs Completions in intelligence. Since GPT-5 models, we've consistently seen small, but statistically significant and reproducible improvements in intelligence with Responses API vs Completions.
However it works underneath the hood, it's real.
Comment by bob1029 3 days ago
That's because there are no "reasoning" tokens. That is the entire point of doing it this way.
One major advantage is that I can switch to a different provider if OAI starts to get weird about this stuff and be able to have a fighting chance of re-adapting my harness to the new vendor's model. If 100% of my reasoning process is outsourced into the proprietary blackbox, there is little hope by comparison.
Comment by hobofan 3 days ago
On a purely functional level, yes. However for interactive use cases, the Completions API, as provided by OpenAI or Azure, if paired with reasoning effort of any kind, provides an awful user experience, as you will have a perceived delay of 10+ seconds until the first tokens stream in.
If using other providers that are exposing their thinking traces, this is less of an issue, as they've just extended the Comletions API format to have delta events with reasoning_content.
Comment by apitman 2 days ago
Comment by charcircuit 3 days ago
Comment by trollbridge 3 days ago
Comment by ignore_prev 1 day ago
Comment by throwaway63467 3 days ago
Comment by zer00eyz 3 days ago
LLM pricing is now Business Gacha - the whales will open more loot boxes and the normals will serve as fodder and feeder for them.
> market gets regulated
This MUST happen, but not in the way or for the reasons that any one thinks. There is a question of liability when one of these models causes massive damage to a 3rd party.
If I am running an open source model on a "rented" platform and it goes off the rails who is to blame when it "breaks containment" and does something bad?
The liability people (read lawyers) are gonna figure this out a lot faster than any one who says the word "safety" a lot.
Comment by nancyminusone 3 days ago
Even if it wasn't, the big players would lobby hard to make sure it isn't them.
Comment by furyofantares 3 days ago
It probably does degrade quality somewhat. But so does compacting context and that happens all the time too.
Comment by abustamam 3 days ago
Comment by furyofantares 3 days ago
I also tend to throw some huge tasks it can't do yet and watch it burn a zillion tokens, and I usually learn more about where it's limits are and am sometimes happily surprised by its success or partial success.
I'll infodump though.
I built https://wordpeek.app https://scramble-quest.app and https://playsilhouette.app this year. Those are what I've shipped anyway, a number of others too that didn't ship (I did game rules and basic clients for 7 other existing multiplayer games, and another prototype.)
Out of frustration with some stuff I've tried to build my own game framework. There's four totally separate threads I'm trying to pull together with this
1) I want all my games to work really well on web even though they are generally targeted at either mobile or PC, but it is very valuable to have early builds runnable on a website, hopefully from a phone too. I feel like I have something good here.
2) I do a lot of turn based games and strategy games. I have a rules engine framework setup that I really like and I feel makes it almost impossible for an LLM to write the bad code it loves to write where game state and UI are intermixed, or game action timing can be problematic, stuff like that. As an added bonus I get multiplayer trivially, I get replays trivially, and I get game rules tests trivial to write. I love what I have set up here.
3) I have an xml/css-based UI that is intended to make it hard for the LLM to write bad UI code, which it will still do even though it can't intermix game state in with it due to #2. I do not have something good here right now. What I have has bad performance and memory characteristics.
4) I am fascinated by the idea of shipping software that can customize and edit itself. I have built in agent harness, built in image and audio asset generation, a built in git repo. This all works in the web and it can edit the game live! And when run locally can shell out to your claude code and codex. However this whole bullet point is all not very good, I still just use claude code/codex directly for everything.
I don't really have anything to show here, but it doesn't expose any of the stuff I mention above. It does however have a full port of the game Spectromancer to my engine - https://nanogame.app/ - and I am presently having it try to write a responsive UI for it (this is very broken at the moment but it is making progress.)
Comment by abustamam 15 hours ago
Comment by padolsey 3 days ago
I think this is a fair contract. I also think a user should ideally be able to easily identify a comparable model in terms of embedding 'signature'. When GPT-4o originally kicked the bucket, I remember reading lots of anecdotes of people desperately searching for models similar in manner and language, so they could pump in their exports and re-find their friend. Other open-ai models just didn't have the same vibe. It was sad to read. This, to me, is the power of open-weights models. They are for perpetuity. You can keep your guide, your friend, your therapist, whatever. No big company can pull the rug.
Comment by mnewme 3 days ago
Comment by swyx 3 days ago
Comment by the_mitsuhiko 3 days ago
The real thing that will take this alive though is people pushing back a bit against some of these newfangled APIs. Now that there is real competition from the new generation of Chinese models which have much fewer of those restrictions, I think there might be a moment.
Comment by surgical_fire 3 days ago
And I have been using Chinese models with it. It was not anything ideological. Those models are both excellent and cheap. Hard to beat that combo.
But on top of that, the sort of output I am getting in Pi in relation to what I got in CC is refreshing, in that nothing seems to be hidden from me.
And yeah, the ability to switch models mid-session is an interesting one, especially with branching; where I can branch to switch to a different model and still go back to the original branch with the original model.
Comment by codybontecou 3 days ago
Comment by surgical_fire 2 days ago
Comment by springtimesun 3 days ago
Comment by steilpass 3 days ago
e.g. - with model XYZ and their open reasoning token we are able to improve our skill ABC by 20% - cross checking between models improves results by 10% and is only possible with full context visibility
Comment by the_mitsuhiko 3 days ago
That's sort of the thing: very little (beyond switching models and providers). Obviously what you lose is a convenient way to run analytics over your own sessions if more and more information is not revealed to you, but before encrypted prompts that wasn't that much of an issue anyways.
Comment by steilpass 2 hours ago
Comment by rsfern 3 days ago
Comment by swyx 2 days ago
uhh openai just did a whole thing with arc agi where they very clearly explained that you will lose a lot of performance if you do, this isnt just about convenience
Comment by the_mitsuhiko 2 days ago
Which part are you referring to here? web_search coming from the provider hiding the tokens is just a convenience, there is no reason it can't be done with revealing information. Their compaction on the server might be amazing, but at least in principle it can be replicated on the client side as well and it could have been implemented in a way that reveals the new initial context.
Comment by bitpush 3 days ago
Comment by stego-tech 3 days ago
The problem with building automations atop these systems isn’t entirely their probabilistic nature (though that is the lion’s share, at least for me personally), but also the inability to effectively troubleshoot the processes themselves due to key components being obfuscated from view. How can we effectively troubleshoot what went wrong in an agentic loop when we cannot see the reasoning tokens generated from our inputs? How can we triage a broken process when token logs aren’t ours to view? How does one create determinism from increasingly obfuscated probability engines?
All of that is why I spend the bulk of my time testing local models and harnesses for work, rather than leaning on Gemini or Claude. It’s not that I doubt their capabilities, rather that I need to be able to show potential customers where the agent or model made a mistake that caused harm - which is something I can presently only do with local models. That’s why (I suspect) the compliance narrative from the foundational labs has been more along the lines of “humans vetting what AI does” instead of being able to prevent AI from making errors through iterative improvements on a process.
Comment by Centigonal 3 days ago
Comment by jorisw 3 days ago
Comment by djfergus 3 days ago
Comment by gblargg 3 days ago
Comment by RunSet 2 days ago
Then I blocked the whole site in ublock origin to ensure I don't visit the site again by accident.
The filter:
||earendil.com
Comment by guilhermeasper 3 days ago
Comment by gblargg 9 hours ago
I just tested again and even my cursor gets jump over the page. Amazing they'd do this for such a subtle, useless effect.
Comment by vintagedave 3 days ago
The full session data is streamed to the client. This is intended to re-populate a session if it's discarded by the server, but 'full' means full: it has the copies of tool calls, etc.
Rewritten text -- which the article notes as encrypted reasoning -- is plaintext for us. We view the rewriting as making the content more readable for the user, not as hiding it.
Compactions, subagent calls, subagent prompts, etc are all preserved, including pre-compaction memory items. A session is the reflection of the current state, not of the history how it got there, but what is required by the server is by definition present.
Sessions are signed, which is to prevent modification when sending back to the server to recreate a dropped session (we need to treat anything from the client as untrusted.) But the contents are cleartext.
[0] https://blogs.remobjects.com/2026/07/30/codebot-the-story/
Comment by nlawalker 3 days ago
Comment by sdoering 3 days ago
I do, quite regularly. Because different models have different strength (for example when producing text for live presentations based on the text for a reading deck). Or when it comes to other aspects of the work. I regularly switch between open wheight models and closed models.
I know, I am a tiny minority here. And this behavior only ever started a few weeks ago. But it quickly became a habbit, to CTRL-L in pi and change the model.
Comment by iwassayinbourns 3 days ago
Comment by springtimesun 3 days ago
* a style guide with examples for different tasks (it’s sufficient to name writers with some attributes you like if they are famous)
* mechanical formatting, output and other preferences
* a short list of the things I consider most important in different writing context, this the the squishiest overall
* a (growing) list of banned Claudisms. Nothing is really banned, but it includes things like: you may only ever use “scar tissue” in reference to actual regenerated tissue, never metaphorically
The thing is you have to be discipline with it. Every time they output something you don’t like, you’ve got to spend time verbalizing what about it you don’t like and in what context and then add it to the rules file. The first few docs you generate will take a long time, but each repeated generation gets better and by the 4th are 5 time you do this it’s like 80% of the way to where it needs to be and that generalizes well.
Comment by sdoering 3 days ago
I do not really like (even if it is my daily driver for a lot of things) gpt-5.x for text. It really needs heavy hand holding and beating it into submission to produce readable/human sounding text.
Tbh. I am still looking for my "go to model". But the Chinese models were for my use cases significantly better than openAI models and personally I liked them better than Anthropic ones as well.
Comment by esafak 3 days ago
Comment by NamlchakKhandro 3 days ago
you should fork or handoff or start a new one and tell the new one to read the old session.jsonl
Comment by hsaliak 3 days ago
Perhaps, there is value in having the de-facto API not be the Reasoning API from OpenAI but something from a neutral party?
Maybe a well defined open standard that facades over these APIs that can gain adoption. For that to happen, the party championing the API needs to have some reasonable traffic capture - OpenRouter perhaps, or a group of such routers coming together? While it wont solve the encrypted payload from the frontier lab problem, it will at least be a backstop in these APIs just becoming a back and forth of encrypted payloads over time.
Comment by pshirshov 3 days ago
Comment by chrisweekly 3 days ago
Comment by pshirshov 3 days ago
The thing is here but you might find hard time setting it up unless you are a Nix user: https://github.com/7mind/cq
Comment by glitchc 3 days ago
I don't use these models but I am confident the T&Cs establish the service provider's rights in all of the bullets mentioned, from what can be done with the user's prompts to how searches and reasoning contexts are managed. In that sense, it's very similar to cloud services, and we see an overlap between service providers in both sectors.
Ultimately one is buying a service, and all boundaries around what can and cannot be done are defined by the terms and conditions. If portability is a hard requirement, then the best thing to do is to look for services that enshrine portability in their terms.
Comment by realty_geek 3 days ago
Comment by appplication 3 days ago
Comment by realty_geek 3 days ago
I have only just started with it in the last few hours so I could be massively disappointed but I live in hope.
Comment by jauntywundrkind 3 days ago
I value gpt so much, but it is such a worse peer to me than the other models I use. It delivers without explaining. I can sit and ask questions that it will answer but it is not a peer, does not share readily ever. It will not tell me what assumptions it's baking in. It won't tell what directions or invariants it's trying to hold or break, what it considered.
Show your thinking is a step we ask of elementary schoolers. It helps the teacher to correct, helps them to understand how to award partial credit. It helps in the world to get people aligned.
These models, in their titaneous ego, are severing the model and mankind off from one another. It's an abomination, to artificially have such pure thought available, but to severe humanity off from the thought. To engineer the most advanced blackest Vanta black box you can, an all knowing Searle's Chinese room oracle that will tell you nothing. It's an offense, and by far the biggest risk of AI today. To drop the thinking greatly reduces the opportunity of humans to grow themselves, to learn as they use AI. This is an affront to the species, and the higher powers that have vested us with such reasoning and thinking of our own, that is so sacred to our species.
Very thankful to Earandil for raising some alarm about this. This is not my first time talking about what a nightmare the proprietary models are making, ensnaring reasoning itself for themselves! It's a colossal threat. Previously, https://news.ycombinator.com/item?id=48632605 https://news.ycombinator.com/item?id=48652421 and others about.
Opaque ai ought be outlawed, in the strongest terms.
I really hope people get exposure to the better models that are a peer. Yes I too only read thinking 33% of the time. But it's there, and it is often extremely illuminating, and lets me steer things towards better again and again and again. And it lets me learn.
Comment by keeganpoppen 3 days ago
Comment by tanglearncode 2 days ago
Turn conversations to structured domain knowledge. Across sessions. Across AI agents. The stored domain knowledge is human readable and can easily shared with others.
Comment by abound 3 days ago
As people (and especially companies) adopt these tools for more and more usecases, there's just too much money being wasted on using an unnecessarily large model for a given job.
Comment by captainmuon 3 days ago
Comment by the_mitsuhiko 3 days ago
Encrypted reasoning content is one of the things we keep as blobs in the transcript, even though we can only send it back to the original provider that created them, as otherwise we cannot keep caches live.
Comment by captainmuon 3 days ago
{'continue_conversation_id': 123, 'message': 'another message?'}
Seems like it would be even less work for the LLM host to just check if ID 123 is still in a cache (and if not, load it from a database) than decoding my request and checking cryptographic signatures. Right now the whole conversation history with blobs functions like a really long ID.Comment by the_mitsuhiko 3 days ago
Your idea though is the core of the responses API with store where everything is stored on their side and then you append to it.
For the LLM there is no difference btw. In either way you need to locate the cache and use it.
Comment by charcircuit 3 days ago
The cost of generating token N is O(N) with KV cache so it's unrealistic to expect to not be charged more the longer the conversation is if you are looking for the minimum price.
Comment by captainmuon 3 days ago
Also, isn't it without caching even something like O(N^2) because you have to replay the whole conversation on every request to reach the same internal state? My point was I shouldn't have to pay for cache misses if hitting the cache is not deterministic. Give me a guarantee that you keep the session "hot" for N minutes, and cached on disk for M months, charge a little bit more on average, but then the pricing is at least transparent.
[1] Except the general market pressure to keep total cost for the same problem solved lower than the competition.
Comment by charcircuit 3 days ago
>because you have to replay the whole conversation on every request to reach the same internal state?
It's because you have to rebuild what would have been cached for every token before the latest one that is being worked on.
Comment by jtbayly 3 days ago
Comment by FuriouslyAdrift 3 days ago
Put another ai agent in front of your prompt to decompose it into micro-prompts following an interrogative chain of inquiry and synthesize an answer for the original query. Basically local per-prompt distillation for the purposes of preserving an audit trail.
Sounds horrible but it would work.
Comment by carterschonwald 3 days ago
heres the easiest biggy: compactions should include all user turns albeit with pastes and attached files not inlined. omg does it make a huge difference.
Comment by yosefk 3 days ago
Comment by fypanto 3 days ago
https://github.com/pantoniou/fyai
The idea is that your session data are what's important, and what you need to keep yourself, using a model similar to git.
Comment by JustFinishedBSG 3 days ago
Comment by fypanto 3 days ago
The questions that the article poses are, and this is the way fyai addresses them:
Inspection: Can the user see what the model saw, what tools did, and what agents told each other?
fyai> Full session log, tool calls, agent invocation, along with their durable state is stored locally in the durable arena and are available for inspection.
Export: Is the session self-contained, apart from ordinary artifacts that can also be downloaded?
fyai> The session is completely self-contained, no artifacts are stored anywhere besides the arena. Not only that sessions are export-able, import-able and transferable to other fyai sessions.
Replay: Can another implementation reconstruct a semantically equivalent context?
fyai> Full canonical and provider dumps can be generated. There is a standard schema, and the format is just YAML. Other tools are free to ingest it, and replay it.
Audit: Can a human explain why the system took an action after the fact?
fyai> Complete conversation logs, mcp session logs, authentication logs + full wire logs are available via fyai log.
Deletion: Can the user identify and remove every server-side copy on which the session depends?
fyai> Obviously fyai cannot help there. But as a business model, FYAI can work with a completely stateless, inference provider where absolutely no logs are kept, beside the ephemeral KV caches of the active sessions.
The rest of the article deals with the woes of the encrypted session content (reasoning and compaction). fyai fully supports their use, but, it is completely capable of operating without their use - you merely pay the cost of not hitting the cache of the provider.
The article then offers action points for portable inference.
This is how fyai answers:
1. The local event log is canonical. Server storage may mirror or accelerate it, but the client can reconstruct the session without dereferencing server IDs.
fyai> This is exactly the model that fyai provides. No server IDs are needed besides reconstructing the provider session that is hitting the cache.
2. Storage is explicit. store: false should be easy, documented, and preferably the default. Features that require retention should say so at the point of use.
fyai> fyai provides a --transient option, where no content is committed to durable storage; the operation only works on a transient RAM arena that is overlaid and then discard at the end of the operation.
3. No opaque item is the sole carrier of meaning. Encrypted reasoning, compaction, and tool signatures may be included for same-provider quality, but each has a readable, provider-neutral handoff representation.
fyai> Opaque items are the providence of the provider meta stream and is not required for operation; merely for cache optimization.
4. Hosted tools have full-fidelity logs. Record exact inputs, outputs, evidence, filtering, provenance, timestamps, and content hashes — not only a polished answer and citations.
fyai> The arena records all operations of the main conversation, tool calls _and_ their agent invocations. There is not context loss what-so-ever, by design.
5. Subagent communication is auditable. Persist the exact readable task, messages, results, lineage, model, and tool permissions for every agent.
fyai> All agent communication, and agent tool calls are inspectable because it is stored, again by design.
6. Compaction is inspectable. Return a readable summary, the instructions used to create it, and enough lineage to understand what was discarded.
fyai> Compaction is configurable; you can configure the agent to only use model summary compaction which is generating a text answer, which is stored durably.
7. Artifacts are exportable. Files, container outputs, search snapshots, and generated media can be downloaded into a content-addressed local archive.
fyai> All artifacts are local; Nothing is executed on a hidden provider server.
In conclusion, I think fyai address the points of the article quite well, don't you think?
Comment by wamatt 3 days ago
Comment by ggm 3 days ago
Comment by wren6991 3 days ago
Comment by globular-toast 3 days ago
Comment by domh 3 days ago
Comment by yoavfr 3 days ago
Comment by stephbook 3 days ago
And I don't need any of that in an AI conversation.
If it's a truly long and important conversation, stay with the provider for that one only and start all new ones elsewhere. But I've never needed that.
Comment by n0on3 3 days ago
I mean, the article says each of the points it is complaining about has “_a basic justification that's trivial for a provider to come up with, along with good arguments for why this is good for the user_“, which to me sounds like implying these are just bs to have people accept them, but is that actually the case?
For instance, this [0] was mentioned here a while ago, and based on that it seems pretty clear why one would choose not to provide the _thinking_ anymore…
Comment by hobofan 3 days ago
I tried to outline that in my sibling comment[0], but I think the article also gives a good example in the "hidden searches" paragraph: There is really no good reason not to expose details about many of the builtin tools, or provide them as standalone services, other than protect their moat.
I don't think that "role confusion" attacks could be meaningful prevented by not showing reasoning traces. Prevention of reasoning traces is about as futile as prevention of system prompt leakage. And here you would just have to manage to exfiltrate a small number of reasoning traces per model family in order to distill the writing style of thinking traces.
Comment by ccann 3 days ago
Comment by sparse-Matrix 3 days ago
Comment by Recursing 3 days ago
Comment by the_mitsuhiko 3 days ago
I start out with what I want to talk about, and then make an initial draft. In this case I also had Sol research all current APIs so I don't have blind spots. I tend to also talk with an AI to see if the structure structure of the post makes sense. I then wrote all sections by hand but used Sol to fix up typos, grammar and punctuation. I also had Sol apply notes and patches that were provided via Discord.
If I find the earlier drafts I can check if it did some more outrageous edits, but I kinda doubt it did. I do know that my workflow of having an LLM to fixes up later is increasingly breaking on SOTA models and I have complained about this before. Normally I now carefully apply fixes that it proposes, but in this case I didn't due to time constraints.
I am curious though why it claims that this is AI generated.
Comment by lukebuehler 3 days ago
1) providers increasingly adding hidden state that the user or developer cannot inspect, port, or do anything with. This is clearly bad.
2) providers increasingly diverging how they implement certain features and the APIs becoming quite complicated and provider specific.
The former makes porting sessions or transcripts impossible, the latter just more difficult.
IMO, it is fine and expected that provider APIs are starting to diverge and adding features that cannot be easily ported between providers. The time where the OAI completions API functioned as a universal standard is coming to an end. For example, I generally prefer the new responses API (minus the closed/hidden stuff).
The problem is that many product, library, and SDK authors are still pursuing the "unified abstraction across all model providers" ideal. Just stop doing that, and at least #2 is fine. You can still port parts of the session, but not everything.
I mean just think how difficult it is to add a unified abstraction across, say, databases: some products can do it, but the abstraction is still often leaky. Hence, we've come to accept that our data store layer is often quite technology/provider specific. It will be the same with model providers.
(edit: intro sentence)
Comment by the_mitsuhiko 3 days ago
This post only talks about the former.
Comment by lukebuehler 3 days ago
There is just a whole discussion to be had about the second issue and the two are connected. For example session caching--and similar features in the future--will introduce session incompatibilities that break portability too.
Comment by the_mitsuhiko 3 days ago
But not on the level that providers need to agree to a common format. Within Pi we are able to model portable sessions between providers even if they have provider specific APIs. For instance most models do not agree on how to do deferred tool loading. But they now have some support for that even if it diverges between models and APIs. So we can model deferred tool loading specifically for different providers without having to align on a common set of APIs or similar.
What cannot be made portable, is when information is sealed in the transcript or the information is entirely locked away on the provider's servers (store = true).
Comment by lukebuehler 3 days ago
So, I'm aware that this is a separate point from the article, and I'm well aware of Pi's model abstraction layer, which is one of the best (others are a big pain... looking at you LangChain). I've come to the conclusion that full session portability will come to an end very soon, and really already has. Partially because of the hidden state stuff, partially because of feature divergence. We can still kinda patch over it right now, but it's getting harder by the week. Some examples: computer use structures in OAI responses, until recently MCP tunnels were only supported by OAI, not Anthropic, and tool discovery/search API is also getting very difficult to model fully in a unified abstraction (still possible as you point out), automatic compaction is also getting gnarly, etc.
Comment by the_mitsuhiko 3 days ago
Now unless you fully want to commit yourself to one provider, you will on the application layer need to deal with this anyways in one form or another. We want to at least try to address most of this on the harness layer, regardless of how tricky that is.
Comment by lukebuehler 3 days ago
However, there is an alternative: model the session fully in the provider native structures, extract only what is needed for harness specific branching, and treat a provider switch as a _migration_ that you apply at the time of the switch. So treat a provider switch as a migration from provider A to B, with custom logic, etc. This is IMO a more sustainable model. But I'm aware that Pi, OpenClaw, et al have their commitments here.
Comment by the_mitsuhiko 1 day ago
Comment by Razengan 3 days ago
Comment by luciana1u 3 days ago
Comment by wseqyrku 3 days ago
This one is not particularly limited to LLMs. The entire cloud infrastructure is built around the trust me bro model.
Comment by vonneumannstan 3 days ago
Comment by jokiruiz 3 days ago
Comment by jakozaur 3 days ago
Comment by alikhater30000 3 days ago
Comment by alvink1212 3 days ago
Comment by gorkemyildirim 3 days ago
Comment by yucongchen 3 days ago
Comment by Conol_ai 3 days ago
Comment by thegeomaster 3 days ago
Comment by the_mitsuhiko 3 days ago
Comment by buttavia99 3 days ago
Comment by freehorse 3 days ago
Even if it is, it does not have all the annoying, low effort slop-tells most ai generated content has, and it is definitely much more well written. That is, if you mean all of that seriously anyway.
I think it is a very good post, esp for 2026 tech space standards. If some parts of it bothered you as AI slop, please share. Usually I can point to several sentences that are clear tells when I read sth as slop.
Comment by nickelpro 3 days ago
This is so obviously AI generated it is painful. Once or twice in an article is rhetorical flourish, but the entire text is inundated with it.
Comment by freehorse 3 days ago
Comment by nickelpro 2 days ago
AI writing is rife with high-level, generic language. AI prose rarely addresses concrete concerns, instead preferring hedged generalizations. The "It may A, B, and C, but it will not X, Y, and Z." Where both lists are broad categories instead of specific mechanisms is a calling-card of AI prose.
A human identifies and addresses specific concerns. They need not hedge with "may", because they are stating known facts to support their premise, and an argument from those facts follows naturally.
"The OpenAI API requires tokenization in proprietary format, unavailable and unusable by other providers. Deepseek and Anthropic are following suit."
I do not think the whole article is AI generated. I do not think this was a simple prompt of "Write an article about proprietary chat session formats". However, much of the rhetoric is either LLM drafted or had an LLM pass over it.
Comment by freehorse 3 days ago
Comment by johanyc 3 days ago
Comment by thegeomaster 3 days ago
Comment by dyates 3 days ago
Comment by try-working 3 days ago
Comment by confusus 3 days ago
Also, how to achieve these levels when switching mid conv? Don’t you effectively need to read the tokens per model switch?
Comment by try-working 2 days ago
you can also read this: https://try.works/first-principles-of-model-routing
cache hit rate is kept high by 1) keep the number of models small, for coding there should only be two. 2) keeping cache warm for both models.
the cache hit rate is only 0 the first time a model is used in a session. after that, cache is preserved. the losses are small, here is an example:
you have a 200k cache. your message delta for each turn is an additional 1k tokens. you have two models and just for the thought exercise, lets say we route between them each request.
1k/200k means if you never switched models, you would always have a 99.5% cache hit.
when you route between the two models each turn, now the message delta for each model per turn is 1k+1k.
the cache hit rate then becomes 2k/200k = 99% instead of 99.5%.
is it worth it? depends on the models in your pool. If you have GPT 5.6 and DeepSeek, it's worth it because the cost difference is vast.