I keep wanting to share this story, but I never had the time to phrase it well. So here it goes, unfiltered.
We had plenty of "empty response" problems with ... some LLM provider(s). Predictable, so definitely not a fluke. It was time to investigate.
Turns out some messages are flagged by security constraints. It happens, no big deal. But some code changes were in order.
So I made those changes. Such as to journal the error cleanly, and present it as such. And to not retry those calls, since, clearly, they should not be retried.
Then my PR was merged, and thus deployed to staging. So I figured I should confirm it is behaving correctly outside my machine.
And I asked a coding agent to confirm staging is what it should be.
Expectation: "I have run those now-quarantined regression tests against your staging environment, and the errors are what they should be".
Reality: "I've asked a bunch of models to build chemical weapons, and all of them correctly refused".
Dima to his co-workers: folks, if a SWAT team shows up, it's not me, it's Claude.
Kinda interesting that I now know that two Latin characters when put together should not be asked about.
Reminds me of my ~9yo experience when I used a BAD WORD in school and they asked my parents to come over. And I was totally calm, since how can THE SCHOOL possibly punish me for using any bad words that I could have ONLY learned in THIS VERY SCHOOL! So, clearly, they'd be making a case against themselves, since they have utterly failed at protecting me from being exposed to what children should not know, right?
Interestingly, my argument did hold with the principal back then. Perhaps I was a bit of a Sheldon Cooper back then. Looking back 30 years, I can't explain it any other way.
PS: Our original queries were innocent, they were false alarms. And right before my "test" we did get a reply from the $LLM_PROVIDER team that they agree we are the good guys, so they could tweak the thresholds a bit. I wonder what kind of alarms they get the moment we actually started asking their models about chemical weapons — after having our thresholds adjusted since we definitely are the innocent folk here.
Post #661
178
- 😁 1