Dolt and Doltgres were using a secondary index for GROUP BY queries even when a full table scan would be faster. By analyzing EXPLAIN output and flame graphs, the team discovered that when a filter selects more than ~50% of rows, the secondary index lookup overhead outweighs its benefits. They added heuristics to the query coster: secondary indexes are only preferred when they select fewer than 25% of rows, while primary keys and covering indexes are always preferred. Using existing statistics histograms to estimate filter selectivity, this reduced groupby_scan latency by 57% on Dolt (144ms → 62ms) and 44% on Doltgres (147ms → 83ms), making Dolt faster than MySQL on this benchmark.
Post #2300
14.6K

No Index GroupBy Optimization
- ❤ 8
- 👍 4
- 👨💻 1