I did two very long experiments with GPT-5.5 xhigh fast, one failed completely, another one succeeded wondefully.
I had a task 2 years ago and it was very hard task - i needed to boot android emulators as EC2 instances or as some VM, but i needed it to be safe and contrained, not some forked QEMU with disabled security. There is still no solution for this - you have to build something yourself, i have asked GPT to boot android in Apple Hypervisor (write swift app to use this api!) and show me screenshot of it and give me a web ui to control it remotely. It implemented it perfectly, it took vanila AOSP build, created a 30+ patch files, build a dockerfile with all scripts to build an image and then booted it on mac. Then i asked to use vanila QEMU and it succeeded too. Then i asked it to benchmark GPU against host, it succeeded too. Then i asked it to make it secure, thats where my token limit ended, but i assume it would be still successful.
Second task is to build an AI runtime. I need a some unified way to run agents that WONT require me writing code, but i still can configure system nicely. That would work with API and with Subscriptions, that would have different sandboxing mechanisms, etc etc. This task failed. It failed because it did work on non-latest main (i forgot to pull) and it built some very very simplified version and completely wrong. I have asked then to pull the main and thats where everything become a huge mess.
So far i am impressed, didn do similar test for Opus tho, curious how it will perform.
Post #638
561
- 🔥 6