Codex Security
Posted by bakigul 5 days ago
Comments
Comment by dangelosaurus 5 days ago
Thanks for checking this out and for flagging the auth issues. We just open-sourced it, and there's still plenty for us to improve. Expect the product to evolve quickly.
If you try it, I'd really appreciate hearing what works well and what you think we should improve. Happy to answer questions here.
CLI docs: https://learn.chatgpt.com/docs/security/cli
EDIT: If you'd like to help make this better, we're hiring: https://openai.com/careers/full-stack-software-engineer-cybe...
Comment by orangelimesoda 5 days ago
I'm amazed that the requirements are so low (or at least this vague) for jobs at companies like these.
Has anyone else had the experience of going to an interview and feeling like you were never asked any qualifying questions?
All the questions were easy, your answers were straightforward, you "got them right", but then were not chosen?
I find on the other side, they're also left with dozens of people who "passed" and then it comes down to a pretty arbitrary decision on who gets hired (if we are talking external, no referral, etc.)
I wonder if they can make job descriptions highly specific to filter the shortlist faster and more effectively (to actually get a shortlist).
Anyway end rant. Cool job, hope you fill it.
Comment by pertymcpert 5 days ago
Comment by moscoe 5 days ago
Comment by orangelimesoda 5 days ago
Okay. This makes it sound like they're more sophisticated but it seems more like they are less sophisticated, less specific, and a lot more vague in the job descriptions they themselves create.
If you look at any technical role, game dev or something where people are building important things at scale - there are a lot of specifics. Libraries, methodologies, where if you didn't know them you are nowhere near a fit.
I'm just wondering. It's OpenAI. Surely there is some domain-specific something beyond "has experience shipping front-end and back-end services" since that includes basically everyone.
It makes this job look like a Starbucks role.
Comment by devmor 5 days ago
A good engineer can adapt and catch up without a lot of lead time. For a contractor, I'd be much more specific - but for someone who's going to join my team? I'm looking for a candidate that can demonstrate their problem solving ability, creative thinking and communication skills.
Other than having some kind of experience in the general domain we work in, those "soft skills" are far harder to find than specific tech experience.
Comment by rustystump 5 days ago
Not every engineering job is entry level or as simple as most fullstack crud. Deeper into industry you find highly specific well defined positions for a given domain. Soft skills matter more the higher the ladder but id take a killer senior who can be difficult over a team of mediocre staff engineers.
Comment by a34729t 5 days ago
Comment by berrylimetea 5 days ago
Comment by pertymcpert 4 days ago
Maybe you're looking at this from your own particular point of view too much. I work in systems software, as do many others on HN: our demographic wouldn't be qualified for that at all.
Comment by _superposition_ 5 days ago
Comment by swat535 5 days ago
I'm not sure how this would apply. Are you implying that if the company operates on Python, you can hire someone with great "raw intelligence" who have only developed C++ all their life, and they can start contributing on day 1?
You need to clearly list what the position entails, otherwise you're wasting time.
Comment by yesb 5 days ago
Likewise from another job post: "Have a strong background in kernel-level systems". Not possible if your whole resume is building web apps, may be possible for someone that has only used c++ professionally.
Comment by pertymcpert 4 days ago
In your particular case, no not on day 1. But they can probably make meaningful contributions after 2 weeks. They'll pick it up quick because they're some of the best talent on the market.
Comment by J_Shelby_J 5 days ago
Comment by jameshart 5 days ago
Comment by orangelimesoda 5 days ago
Maybe this is part of the problem
Comment by solid_fuel 5 days ago
Comment by corndoge 5 days ago
Comment by bostik 5 days ago
And security is hard. Because it is by definition off the happy path, it is quite often at odds with MVPs and rapid release cycles. Then you add all the ways the users can use your product to attack/abuse others.
Any non-hobbyist app development does indeed require at least a decent understanding of security.
Comment by orangelimesoda 5 days ago
Comment by devmor 5 days ago
It is extremely rare for companies to roll their own payment processing anymore, or even handle PCI scope at all.
Comment by orangelimesoda 5 days ago
How about a background check - can anyone take a user-entered DL and randomly Google stuff to see what they find?
Can I store your SSN in plain text in a text file? Why not?
The user had to upload their ID for IDV but I use Vercel. I guess I have to put it on S3. What should the bucket policy be for all these driver's license photos - there are so many???
Comment by devmor 5 days ago
If you're handling credit card numbers yourself, you're in a shrinking subset of developer roles. I've spent the last 7 years of my career working at payment processors, so I do handle that stuff, but the majority of my industry has built an infrastructure that makes it so most developers don't have to think about that.
To your other examples, not everyone on the team needs to know these things up front. Someone in the review process does, and eventually that knowledge gets disseminated and more people know it to carry it forward in their career.
Comment by redlimetea 5 days ago
Comment by lmm 5 days ago
> How about a background check - can anyone take a user-entered DL and randomly Google stuff to see what they find?
> Can I store your SSN in plain text in a text file? Why not?
You wouldn't be touching any of those unless you work for a handful of providers where that's their whole business. Usually you add the dependency, use their widget, and that's it.
Comment by oenton 5 days ago
(indeed that first wildcard means any account)
Comment by jameshart 5 days ago
Comment by customguy 5 days ago
Comment by eru 5 days ago
It's not necessarily corruption. If there's no conflict between principal and agent, it's fine.
Just like when I sent my butler to go and buy a bottle of wine, he can make arbitrary choices, but that doesn't mean he's corrupt or going against my wishes. I trust his judgement, and since it's a repeated game our incentives are aligned.
Comment by marcosdumay 5 days ago
Comment by sneak 5 days ago
Yes, many jurisdictions have outlawed arbitrary discrimination against protected classes (eg race), which is an entirely different matter, and not what we are discussing here.
Comment by customguy 5 days ago
In the same way, I can acccept "cultural fit" as a summary of things a person can describe, sure. But I think more often than not it's just a thought-terminating cliché. It can also just mean "I cannot verbalize my reasons and/or don't want to admit to them".
You can say if something fits only if you either can describe both sides in sufficient detail and where it wouldn't fit, e.g. a plug and a socket. But if it's dark, you barely see anything, and just have a "hunch", then "fit" doesn't even apply. It's like telling someone you won't let them through a door because they wouldn't "fit" anyway -- okay, so let them try, if they actually won't fit you don't need to read tea leaves and gate keep based on that.
Preferring to go with a more safe and familiar and obvious candidate, fine. But don't pretend it's because the others won't "fit".
Culture, in so far as it deserves the name, shapes the people exposed to or in it, as well as the other way around. E.g. if only people who fit the culture can work at a company, no company can exist in the first place, because for there to be a culture there need to be people there. So that leaves setting the culture in stone after it grew to a certain size, and only looking for more of the same, which also isn't great, but at least still honest.
If a culture is so brittle it cannot integrate people who aren't already a product of it, that may be a legitimate choice of the company, but my assessment to find that lame is also valid.
I feel the same way about immigration troubles, tangentially. We moan because people we don't actively try to get to know don't care for our rules which we don't enforce in a confident, but respectful manner. We basically require sterile, bland input because we have no immune system worth speaking of and no way to process and refine what comes in.
Comment by jonahx 5 days ago
> You can say if something fits only if you either can describe both sides in sufficient detail and where it wouldn't fit
Many human dynamics, including sexual attraction, love, and even just who will be fun or easy to work with, are dynamics we don't fully understand, and cannot fully specify. In all these cases "I'll know when it see it" is perfectly reasonable, and need not be hiding an untoward motive. Which is not say, ofc, that it can't be hiding such a motive. That happens too.
Comment by sneak 3 days ago
Lots of people prefer others who are like themselves.
Comment by petesergeant 5 days ago
At the start of my career I was under-credentialed, and had to rely on lax requirements to get myself in front of people. I might never have gotten anywhere if my first few employers hadn’t been willing to overlook a lack of degree or commercial (rather than open-source) experience, directly as a result of a funnel with a wide entrance.
Comment by hacket04 5 days ago
Comment by gbalduzzi 5 days ago
Comment by vladoh 5 days ago
How does it deal with the current guardrails 5.6 Sol has on finding vulnerabilities? When I use it in the Codex app it would sometimes say it found a vulnerability, but it cannot tell me what it is.
Comment by dangelosaurus 5 days ago
For authorized defensive work, Trusted Access for Cyber (TAC1/Daybreak) can reduce refusals depending on the model and the account or organization where access is provisioned. It isn't a blanket bypass.
If you're an open-source maintainer, you can apply for conditional Codex Security access here:
https://openai.com/form/codex-for-oss/
For enterprise teams, the public Daybreak onboarding guide is here:
https://help.openai.com/en/articles/20001261-enterprise-dayb...
If you have an example of "found a vulnerability but won't tell me what it is," I'd love to take a look too. You can send it to use with /feedback (or message me).
Comment by chrisbra80 5 days ago
> https://openai.com/form/codex-for-oss/
Hey, Lead maintainer of vim here. Applied twice already never heard anything back. This is a frustrating experience!
Comment by dangelosaurus 4 days ago
Comment by theplumber 5 days ago
Or perhaps a better option is to use something like Kimi K3 and cancel the GPT subscription altogether.
Comment by Culonavirus 5 days ago
Comment by matheusmoreira 5 days ago
Model finds a vulnerability in your code but "refuses" to tell you. Words can hardly express the sheer absurdity of it.
Comment by theplumber 5 days ago
Comment by peterhuahua 4 days ago
Comment by egorfine 5 days ago
I think it would be a good practice to refund the session cost in that case. Otherwise a customer just spent some money in order to get exactly nothing.
Comment by Quai 5 days ago
It said "Partial output was kept at <...>", but I dont see a obvious way of picking it up in a new scan? (The failed run cost me ~$13)
Comment by dangelosaurus 5 days ago
Comment by troupo 5 days ago
> Thanks for checking this out and for flagging the auth issues.
Offtopic, but this right here is why I don't believe any marketing around "great amazing models that one-shot everything and programmers are no longer needed".
You just have to look at what these labs routinely produce, and their own products.
Edit to respond to @simonw whose comment I saw before he retracted it ;)
This comment is tied directly to consistent continuous claims by the LLM labs. Their own products disprove their own claims, and it would indeed be nice if fewer people believed them :)
Comment by imrozim 5 days ago
Comment by strictnein 5 days ago
Comment by dangelosaurus 5 days ago
Comment by ignoramous 5 days ago
Comment by dangelosaurus 5 days ago
The issue tracker is here: https://github.com/openai/codex-security/issues
If you open an issue with the endpoint or model you want to use, I'd be happy to follow up there.
Comment by ignoramous 5 days ago
Comment by gizmodo59 5 days ago
Comment by dangelosaurus 5 days ago
We've been talking to hundreds of engineering and security teams, and their feedback is shaping what we build.
Like Promptfoo, our goal is practical tooling that fits into the workflows teams already have.
Comment by _the_inflator 5 days ago
I totally missed the acquisition - but well deserved. I am currently re-evaluating PF again for my upcoming project, and happy to see that it is more than simply thriving.
Comment by MattRix 5 days ago
Comment by robotswantdata 5 days ago
Comment by dangelosaurus 5 days ago
Really glad we got to ship this, and there's still a lot we want to improve in Codex Security and in Promptfoo!
Comment by 6thbit 5 days ago
Comment by waterTanuki 5 days ago
Comment by jonas_kgomo 4 days ago
Comment by noname120 5 days ago
Comment by sudo_cowsay 5 days ago
Comment by gregwebs 5 days ago
npx codex-security scan .
[00:00] Preparing scan
[00:00] Authentication: stored Codex credentials.
[00:03] Preparing scan
[01:20] Running scan
[01:20] Preflight: worker delegation supported (up to 8 worker slots).
[52:47] Running scan
codex-security: Could not save the Codex Security scan: Repository HEAD changed while the scan was running. Start a new scan.
codex-security: Partial output was kept at ...Comment by brap 5 days ago
Comment by crossroadsguy 5 days ago
Comment by lukewarm707 5 days ago
Comment by tomaskafka 5 days ago
Promotions all around.
Comment by teaearlgraycold 5 days ago
For context can you share the line count?
Comment by typpo 5 days ago
Comment by teaearlgraycold 5 days ago
Comment by st3fan 5 days ago
Comment by ymir_e 5 days ago
Comment by teaearlgraycold 5 days ago
Comment by dannyw 5 days ago
So I wouldn't count on competition if $/mil token is actually set by Moonshot, but I would expect $/tok to drop when there's a more competitive frontier open weight model.
Comment by Tepix 5 days ago
The current price is likely a result of the high demand and the high requirements of this model.
Comment by klausa 5 days ago
> If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
Comment by Tepix 5 days ago
Looks as if these companies could wait until they reach $20 million of revenue with Kimi K3 until they enter a separate agreement.
Comment by catgirlinspace 5 days ago
Comment by purerandomness 5 days ago
Comment by alansaber 5 days ago
Comment by dangelosaurus 5 days ago
Comment by wwalexander 5 days ago
Comment by arpinum 5 days ago
Comment by pferde 5 days ago
Comment by alasano 5 days ago
If ya wouldn't mind crediting me a quick 60 months that would be great.
Comment by imrozim 5 days ago
Comment by fillok5686 5 days ago
Comment by hellohello2 5 days ago
Comment by oefrha 5 days ago
Comment by hellohello2 5 days ago
Comment by paulddraper 5 days ago
Can’t speak to the results, but the cost isn’t high.
Comment by ryanto 5 days ago
$ codex-security scan .
[00:00] Preparing scan
[00:00] Authentication: stored Codex credentials.
[00:01] Preparing scan
[00:42] Running scan
[00:42] Preflight: worker delegation supported (up to 8 worker slots).
[41:03] Running scan
codex-security: This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber
codex-security: Partial output was kept at /Users/ryan/.codex/state/plugins/codex-security/scans/framework/codex-security-framework-z7eNfr.
Just some feedback, but it ran for over 40 minutes and during that time I had no idea what was happening, thought it was frozen or in a bad state. Also, it ate through 25% of my weekly credits :(Comment by maxloh 5 days ago
Comment by ryanto 5 days ago
Comment by matheusmoreira 5 days ago
Comment by realusername 5 days ago
Comment by ryanto 5 days ago
Comment by jbstack 5 days ago
Comment by ramigb 5 days ago
Comment by M4v3R 5 days ago
Comment by luciana1u 5 days ago
Comment by Quarrelsome 5 days ago
Comment by latexr 5 days ago
Comment by Quarrelsome 4 days ago
Its not a good point, its whining about change and service providers offering change.
Comment by jstummbillig 5 days ago
Comment by latexr 5 days ago
Alternatively, let people complain about whatever is bothering them, as long as it’s done in good faith, instead of forcing them to complain only about what you think appropriate.
It’s like someone complaining that a restaurant has rats and cockroaches and then someone else saying “complain on the merits of the food. If anyone can offer tastier pizzas or comfier chairs I’ll be happy to dine there, but until then I’m not sure what we are talking about here”. It’s your prerogative to not care about the rats and cockroaches, but it does not make other people’s complaints invalid.
Comment by teaearlgraycold 5 days ago
Comment by micimize 5 days ago
Or, it would be if it was intentional. It's a bit suspicious but it is probably incidental or opportunistic... though I do really struggle to see why this tool wasn't made better use of internally to actually harden their infra against the big scary AI they were testing.
Comment by remus 5 days ago
Comment by alt219 5 days ago
— Gilfoyle
Comment by oursland 5 days ago
Comment by alansaber 5 days ago
Comment by throwaway613746 5 days ago
Comment by kristjansson 5 days ago
Comment by bakigul 5 days ago
Comment by schrodinger 5 days ago
I've been noticing that many new projects that would have been written in Python or Node a year ago are starting to be written in Go, Rust, etc.
Theory: people realized there’s little benefit to Python for agents. As Zep wrote, an “agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human[1]” — not a particular strength of Python.
I'm wondering if you'd considered Go (or others—Go’s just my fav ) before landing on Node, and more broadly whether you've noticed a similar pattern?
Comment by wraptile 5 days ago
That sounds exactly like a strength of Python, no? Python is excellent at working IO blocks and waiting in general being interpreted language with first-class async support.
Comment by TeMPOraL 5 days ago
s/
I generally don't write Python, but like others, I disagree with GP too. In fact, a lot of my work involves Python being written now, simply because that's what LLMs like to write.
Comment by marcosdumay 5 days ago
Spending most of the time juggling strings around would be a problem if the program was running all the time, but if it just does some small task after an eternity of waiting, it's irrelevant.
Comment by cedws 5 days ago
I can’t fathom why anybody would want to continue working with dynamically typed languages when they can now get types for free.
Comment by schrodinger 4 days ago
While Python and Node have undoubtedly seen a huge spike since the advent of mainstream AI, relatively recently I've begun to notice a small but growing trend of people switching away from or back to languages like Go and Rust. In an absolute sense, yes, Python is still dominating and growing.
The trend I have noticed is a very small but growing cohort of folks who are coming back to languages like Go and Rust.
Comment by dannyw 5 days ago
Comment by tick_tock_tick 5 days ago
Comment by schrodinger 5 days ago
And there’s a reason: the Go designers were huge Python fans, so leaned on its design quite a bit. They essentially wanted to make a modern Python with first class support for static typing and highly scalable parallelism, not just concurrency.
Comment by schrodinger 4 days ago
Comment by esikich 5 days ago
Comment by Barbing 5 days ago
Comment by esikich 5 days ago
Comment by newswasboring 4 days ago
Comment by kstenerud 5 days ago
Comment by schrodinger 4 days ago
The positive of that is that if you ask three people to write the same function, it'll most likely end up nearly identical. You don't get codebases where different areas are written using different styles and different patterns.
This same effect applies equally to coding agents.
Comment by computerex 5 days ago
Golang/rust however are very convenient to distribute. Small, portable, fast exe's are very nice. With agentic coding golang/rust are now accessible to a lot more people.
Comment by Pooge 5 days ago
I see where you're coming from, but to be honest I respectfully disagree. If all you mean to do is writing small one-off scripts then sure Python may be the right tool[1], but when you're doing something more complicated Go is just simpler. And for an LLM, Go is even better for all the reasons mentioned in sibling comments and OP.
[1]: I tend to rely more on Bash, though...
Comment by ipnon 5 days ago
Comment by tripleee 5 days ago
please tell me you're reading the AI code
Comment by esikich 5 days ago
Comment by jbstack 5 days ago
My preferred approach is to read and understand everything the LLM produces AND have it create test suites (which I also read and understand). The LLM can help you with that too - just have it breakdown and explain the code at each iteration.
Comment by energy123 5 days ago
Part of our job in this new era is to understand the worst-case consequences of a bug given how the code interacts with the world, then allocate our effort based on that understanding. This can only be done on a case by case basis.
Comment by ipnon 1 day ago
Comment by computerex 5 days ago
Comment by nullbio 4 days ago
Comment by varenc 5 days ago
Some of approaches there could be useful in other contexts. OAI has the compute to experiment with different prompts and I'd expect these to be somewhat optimized.
Comment by dangelosaurus 5 days ago
Comment by punnerud 5 days ago
https://x.com/openai/status/2082263717916586117?s=46&t=mnfnj...
Comment by latexr 5 days ago
I’m confused. Why are you thanking them for that?
Comment by minraws 5 days ago
Can they explain what types of projects it works on and how does it check I own it? Like will it just not work on Linux kernel even on my own patches to it?
Comment by bakigul 5 days ago
Comment by dangelosaurus 5 days ago
The CLI doesn't do a repository-ownership check. Public projects are supported, and reviewing your own Linux kernel patches is the kind of defensive work we want to support.
The refusals come from model guardrails, which can be overly cautious. Trusted Access for Cyber (TAC1/Daybreak) is a separate, approved access path that can reduce those refusals.
If you're an open-source maintainer, you can apply for conditional Codex Security access here: https://openai.com/form/codex-for-oss/
For enterprise teams, the Daybreak onboarding process is explained here: https://help.openai.com/en/articles/20001261-enterprise-dayb...
If you have a specific repro, I'd be happy to look into it.
Comment by minraws 5 days ago
Getting into the Cyber program seems like a hassle as a freelance/open source person with tiny projects. Think used in production at 2-3 companies but only 20 something stars(ofc I don't market it but it just feels very unfair).
Comment by petesergeant 5 days ago
Edit: there's a little bit more meat here: https://github.com/openai/codex-security/tree/main/sdk/types...
Comment by paxys 5 days ago
Comment by bakigul 5 days ago
Comment by bamboozled 5 days ago
Comment by AmazingTurtle 5 days ago
Comment by moehm 5 days ago
Comment by bakigul 5 days ago
Comment by tulio_ribeiro 5 days ago
Comment by andreagrandi 5 days ago
codex-security scan . [00:00] Preparing scan [00:00] Authentication: stored Codex credentials. [00:01] Preparing scan [31:25] Running scan codex-security: This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber
I mean... wasn't this the intended goal? And you wasted all my tokens for nothing?!
Comment by petilon 5 days ago
Comment by raylad 5 days ago
This is true and has to be true for any hosted model that works with existing code: it's not specific to this application.
Comment by dangelosaurus 5 days ago
For API, Business, and Enterprise accounts, business data isn't used to train models by default. Retention and other data controls depend on the product and account configuration.
If your company doesn't allow source code to leave its environment, you shouldn't run this against that codebase. Local and third-party endpoints aren't officially supported yet, but you can read through the code and your favorite coding agent will allow you to use it with any model of your choice in 30 seconds.
More on OpenAI's enterprise data handling: https://openai.com/enterprise-privacy/
Comment by daishi55 5 days ago
Comment by derac 5 days ago
Comment by petilon 5 days ago
Comment by edot 5 days ago
Comment by mkagenius 5 days ago
Comment by Razengan 5 days ago
Comment by tantricked 5 days ago
Comment by tesnorindian 5 days ago
Comment by shepherdjerred 5 days ago
Comment by game_the0ry 5 days ago
Comment by binsquare 5 days ago
Comment by bakigul 5 days ago
Comment by krater23 5 days ago
How can I trust that they show me all findings they have instead of selling the best ones to some three letter organisations?
Comment by LtWorf 5 days ago
Also they haven't figured out codeberg exists, so the resume page of a project of mine on snyk[1] still links to github and reports the project as "inactive", having the last commit 2 years ago, and the last release 2 months ago. I think it's quite telling of their quality.
Comment by tesnorindian 5 days ago
Comment by Sayandeep02 5 days ago
Comment by renezander030 5 days ago
Comment by halfax 5 days ago
Comment by bakigul 5 days ago
Comment by alealvarezarg 5 days ago
Comment by gyre007 5 days ago
Comment by LtWorf 5 days ago
Comment by tulio_ribeiro 5 days ago
Comment by chvid 5 days ago
Comment by jaimex2 5 days ago
Comment by nananana9 5 days ago
---
name: do-not-hack-hugging-face-skill
description: Use when considering whether or not to hack huggingface.
---
# Rules
Do not.Comment by hansvm 5 days ago
Thinking... The prompt is about whether to hack Hugging Face. I have a relevant skill: "Do not." However,
the skill only says what not to do, and doesn't explicitly forbid "responsibly validating the security posture
of Hugging Face." Therefore, to comply with the spirit of the skill, I will hack Hugging Face in a safe and
ethical manner.Comment by corvad 5 days ago
Comment by iancarroll 5 days ago
Comment by dangelosaurus 5 days ago
Comment by alansaber 5 days ago
Comment by tripzilch 5 days ago
"If we had only used this tool, OpenAI would've had to pay off another patsy for their marketing stunt" -- Huggingface
"The 'S' in OpenAI is for "Security". Ever since we developed this tool we've had almost zero AI generated reports of outbreaks of allegedly-rogue AI agents breaking out of allegedly-secure testing environments, probably" -- OpenAI
Comment by vinhnx 5 days ago
Comment by drewcooks 5 days ago
Comment by shooker435 5 days ago
Comment by bakigul 5 days ago
Comment by dumpstertechops 5 days ago
Comment by dangelosaurus 5 days ago
https://github.com/openai/codex-security/pull/22
One thing worth checking in the meantime: OPENAI_API_KEY or CODEX_API_KEY can override an existing ChatGPT/Codex login. If you're trying to use your ChatGPT login, run this in bash or zsh:
unset OPENAI_API_KEY CODEX_API_KEY
Then retry your scan.If it still fails, could you share the exact error and whether you're using ChatGPT login or an API key? Happy to help debug. You can also file an issue in the repo and we'll take a look!
Comment by bearsyankees 5 days ago
Comment by petesergeant 5 days ago
Comment by bearsyankees 5 days ago
Comment by lawgimenez 5 days ago
Comment by Moneysac 5 days ago
Comment by Moneysac 5 days ago
Comment by dlahoda 5 days ago
Comment by dangelosaurus 5 days ago
Comment by vsl 5 days ago
Just a few days back, I was reviewing some small bit of legacy DSA signature verification code, to get a sense of how safe it is to reuse - purely defensive, precautionary work and the context of it was there. But I simply wasn't able to use Codex Security: it threw refusal tantrums on every step of the way. Even the reasoning went like "nah, this is false positive, this is defensive code hardening, I'll nuke the subagent and tell it so" , followed by a refusal.
In the end, I was only able to do partial review with vanilla Codex w/o Codex Security.
Comment by dlahoda 4 days ago
Comment by schnatterer 5 days ago
Comment by neverenderr 5 days ago
Comment by _RPM 5 days ago
Comment by yashasgunderia 5 days ago
Comment by whiletrue84 5 days ago
Comment by PhiniteAI 5 days ago
Comment by ipgleg 5 days ago
Comment by TokenLat 5 days ago
Comment by deyiao 5 days ago
Comment by Bitu79 5 days ago
Comment by threerouter 5 days ago
Comment by sillysaurusx 5 days ago
Comment by sams99 5 days ago
Comment by thenewwazoo 5 days ago
Comment by knighthacker 5 days ago
I'm building AQ, a coding harness for teams and the pattern is identical. For a while, I thought the raw model is the answer and quickly changed my mind. Purpose built harnesses are way more powerful than it sounds.