Uber Code Review AI Assistant
Uber continues to share their experience to integrate AI into different parts of development process. This time it's GenAI code review assistant (previously they published about GenAI On-Call Copilot and GenAI Optimizations for Go).
If you tried to make a code review with some GenAI tool you may notice it's not perfect yet: hallucinations, overengineering, noisy suggestions. It left the feeling that it produces more issues and consume more time than a human review process.
That's why Uber engineers created their own review platform.
So let's check what they implemented:
🔸 Define relevant files for analysis: filter out configuration files, generated code, and experimental directories.
🔸 Include PR changes, surrounding functions and class definitions to the LLM context.
🔸 Execute analysis calling number of different AI assistants:
- Standard: detects bugs, exception handling and logic flaws.
- Best Practices: enforces Uber-specific coding conventions and style guides.
- Security: checks application-level security vulnerabilities.
🔸 Execute another prompt to check quality of the previous step, assign a confidence score and merge overlapping suggestions.
🔸 Run a classifier for each generated comment and suppress categories with low developer value.
🔸 Publish result comments on PR.
Authors reported that the whole process takes around 4 minutes and already integrated with all Uber's monorepos: Go, Java, Android, iOS, Typescript, and Python.
One more interesting point, that for code analysis and comment grading 2 different models were used: Claude-4-Sonnet and OpenAI o4-mini-high.
As you can see, more and more AI systems start working in multiple stages, where one AI checks the results of another. This pattern is becoming popular and it shows really good results removing noise and decreasing the number of hallucinations.
#engineering #ai #usecase
Post #206
318
- ❤ 4
- 👍 3