Several days ago, OpenAI introduced Codex, a cloud-based software engineering agent.
Today, I realized that I have access to it and decided to look around.
I connected it to my open-source repo https://github.com/PavelPolyakov/fastify-blipp. When you connect the repository, Codex creates an environment for it and runs some "default" tasks against the repo.
Among the first "default" tasks, which started the execution automatically (smart), there was:
* 🐛 Find an important place, find a bug, and fix it
🧐 My experience:
1️⃣ Codex found something and drafted the code change (🖼 #1 in the comments)
2️⃣ But it was not able to run tests; it shared the error that
cross-env is missing3️⃣ I asked it to fix it
4️⃣ 1+ minute of public thinking and executing bash commands. It suggested changing the test command (🖼 #2 in the comments)
5️⃣ I asked why it removed
cross-env6️⃣ 1+ minute of public thinking and executing bash commands. It said that it did not remove anything and cross-env is in place (it was a lie). It suggested installing dependencies before the test run.
7️⃣ I asked to install dependencies and run tests
8️⃣ 1+ minute... it said that the test run still fails
👩⚖️ Overall observations:
* It's a good start for the product; things feel well thought out
* I like that they stream the model's "thinking," and it's possible to follow it, but you may also ask yourself, WTF are you doing?
* When the model "thinks," it feels like you have no impact (maybe I can get used to it)
* It was also looking for an
AGENTS.md file to read instructions for the coding agent. If I would have it, I guess it would have been beneficial.* I had no expectations and haven’t learned any specific way of working with the engineering agent so far (other than a "simple" chat). So, I consider my experience relevant.
Overall, it still looks like a loop where the model is executed with some context, and that model has access to the "tools". Pretty much like this.
Have you tried it? Feel free to share your experience.