AI SQL Tuner

AI SQL Tuner

Does AI Actually Tune SQL Server? What Our AI SQL Tuning Benchmark Found

Every vendor with an AI feature will tell you it makes your database faster. Very few will tell you how they know, so we built an AI SQL tuning benchmark to find out.

We wanted to answer three questions:

  1. Can AI help with SQL tuning?
  2. Does the model matter, and if so, which one is best?
  3. How much does reasoning effort change the result?

The rule was simple: nothing counts unless the recommendation is actually applied to a real database and the workload is measured again afterward. Either the workload got cheaper and faster and still returned the same answers, or it didn’t.

Here’s what we found.

Executive summary

1. Yes, AI helps — and the effect is not subtle. Every one of the 35 runs on the latest release of AI SQL Tuner Studio cut the work the database had to do, and none returned a wrong result. Typical results:

WorkloadLogical readsTotal workload time
Query tuning (one slow report query)95% less98.5% less
Index tuning (OLTP application workload)79% less41% less
Code review (7 application objects)45% less30% less

Not one of those 209 runs made the workload measurably worse. That is the headline, and it held for all five models.

Bar chart comparing median reduction in logical reads and total workload time across three workloads: query tuning, index tuning, and code review.
Does ai actually tune sql server? What our ai sql tuning benchmark found 4

2. The model matters most on code review. We screened eight models, including all three new releases, and took two finalists through the full benchmark: GPT-5.6 Luna and GPT-6 Sol. On index tuning and query tuning they tie — the differences are smaller than the test rig’s own run-to-run variation. On code review, GPT-5.6 Luna is measurably better, and it held that lead in two separate runs. Both answer in about a minute.  Our pick for most work is GPT-5.6 Luna. GPT-6 Sol earns its place as the more consistent index tuner.

3. Medium effort is the sweet spot. Effort changes how long the model thinks, and on code review it shows. At low effort, code review suffers: GPT-5.6 Luna gave only a minor improvement in three of five runs, and GPT-6 Sol returned wrong results twice. High effort matched medium’s results while taking one and a half to nearly three times as long to answer. Medium is the default in AI SQL Tuner Studio.

The practical takeaway for a DBA: run it, read it, and apply what you agree with. Stay on medium unless you have a reason not to.

Terms worth defining

A few pieces of vocabulary appear throughout. Skip this if they’re already familiar.

  • Logical reads â€” the number of 8 KB pages SQL Server touched to answer a query. It’s our headline metric because it’s deterministic: it doesn’t change based on a warm cache, how busy the box is, or how many CPU cores the optimizer decided to use.
  • Noise floor â€” how much the measurements move on their own when nothing has changed. We ran each workload repeatedly against an untouched database to find out. Any improvement smaller than the noise floor is reported as “no change,” not as a win.
  • Median — the middle result of several runs. We report medians rather than averages so that one spectacular run can’t set the number.
  • Replicate — one model answering the same question again, independently. Five per model per test.
  • p-value â€” roughly, the odds of seeing a difference this big by chance if there were no real difference at all. Small is convincing; 0.08 is not.
  • Reasoning effort â€” low, medium or high. Higher effort means the model thinks longer before answering.
  • Query Store — SQL Server’s built-in history of which queries ran, how often, and what they cost. Index tuning leans on it heavily.
  • Execution plan — the optimizer’s step-by-step recipe for running a query. It shows why a query is slow.

How the AI SQL Tuning benchmark was built

The whole design comes down to one loop, run identically for every model:

  1. Revert the database to a pristine image (a database snapshot, so it takes seconds, not minutes).
  2. Build history — run the workload long enough that Query Store holds a realistic record of it, the way a production database would.
  3. Measure the workload — warm-up passes, then several measured passes.
  4. Tune — hand the AI exactly the same input every model gets, and time how long it takes to answer.
  5. Apply the recommendations, script by script, recording which ones ran and which failed.
  6. Measure the same workload again, then revert and repeat for the next model.

The history step matters because index tuning works from what Query Store says the workload costs, so the benchmark gives it 40 passes of history before any advice is requested. Two more details come from AI SQL Tuner Studio itself: Query Tuner sends the query’s estimated execution plan along with its text and table metadata, and reasoning effort is passed straight to the model, so low, medium and high change how long the model actually thinks.

Some details that make the numbers defensible:

Repeatability. Same physical server, same database image, same workload, same prompts, same statistics, same warm-up. The only two things that varied were the model and the reasoning effort. Within a test, every model answered a pinned input — one captured snapshot of the data the tool collects, replayed identically — so the comparison is between models, not between whatever each happened to see that day.

The noise floor comes first. Before any AI credit is spent, each workload runs repeatedly against an unchanged database to measure how far the numbers drift on their own. Logical reads came in at 0.0%. Total workload time is noisier: 12.4% on index tuning, 10.3% on code review, and 25.5% on query tuning. Every claim below is gated on those floors.

Benefit and cost are measured separately. Indexes make reads cheaper and writes more expensive. Read metrics and write metrics are reported side by side and never summed, so one can’t quietly cancel the other.

Statistics are frozen. Sampled statistics are refreshed once into the pristine image and then left alone, with automatic statistics disabled. A freshly redrawn histogram can flip plan choices, and we don’t want that mistaken for an index’s benefit.

Correctness is a gate, not a metric. Every read query’s results are hashed before and after. A rewrite or code revision that changes the results voids the run — its measurements are discarded, not averaged in.

Measurement uses Extended Events, so CPU and duration are microseconds as the engine saw them. The harness’s own monitoring runs on a separate connection and stays out of Query Store, where it could otherwise look like part of the workload.

The tests

Three workloads:

Index Tuning (OLTP). The shape a real application produces: 4,002 short, parameterized read executions per pass across eight statement types, plus 480 single-row writes that commit. Writes are about 11% by count. Every foreign-key column already has an index, so a model can’t win by recommending the obvious — and the writes charge each new index its maintenance cost.

Query Tuner. One genuinely slow report query: the top answerers of 2020 among high-reputation users, with a correlated subquery per figure. The vote count is the expensive one — the column it filters on has no index, so the optimizer re-scans for every qualifying user. Both kinds of advice have room to show: an index removes the repeated scans, and so does a rewrite that aggregates once and joins. The date predicates are non-sargable on purpose.

Code Review. Seven application objects with problems a fix makes measurably cheaper: a scalar function that can’t be inlined, a cursor, non-sargable date filters, concatenated dynamic SQL, and an OR across a SELECT * view. The OLTP writes run beside it, because reviews recommend indexes too. A review succeeds if the workload got cheaper and faster and every read still returns identical results.

The models

We started with a wide field: GPT-5.4, GPT-5.6 Luna, GPT-5.6 Sol, Claude Opus 4.8 and Claude Opus 5, plus the three new releases, GPT-6 Luna, GPT-6 Sol and Claude Opus 5.5. A screening round narrowed the field to two:

  • GPT-6 Sol was first or tied for first on index tuning and query tuning, answered in under a minute, and cost about a third of GPT-5.6 Sol. It replaces GPT-5.6 Sol as the premium option.
  • GPT-6 Luna was the cheapest model per analysis, but it was measurably weaker on code review than GPT-5.6 Luna. It was also the only one of the three new models whose advice ever made a write-heavy workload slower. GPT-5.6 Luna stays as the standard model.
  • Claude Opus 5.5 matched the others on benefit but took three to four times as long to answer. For the same data it also consumed about 58% more input tokens, which makes every analysis more expensive.
  • The older models were retired, because a newer option beat each of them on speed, cost or code-review reliability.

The two finalists then went through the full benchmark on the latest release of AI SQL Tuner Studio. Every number in the results below comes from those runs.

ModelProviderIn AI SQL Tuner Studio
GPT-5.6 LunaAzure OpenAIDefault model, all editions
GPT-6 SolAzure OpenAIPremium option

Claude Opus 5 ran only the low-effort tests before we stopped testing it: it finished last on every measure that separated the models, and took roughly seven times as long as the fastest model to answer. (Note: Claude Opus 5 was used to develop the benchmarking harness and execute the testing, as well as write this report.)

How many runs, and what each cost

The results rest on 65 measured model runs: 35 on the latest release at medium and high effort, and 30 in a code-review effort comparison at low, medium and high. That’s on top of the noise-floor runs, which involve no model at all, and the screening runs that picked the finalists. Every run included a call to the AI endpoint, a database revert and a full before-and-after measurement.

Results

Question 1 — Can AI help with SQL tuning?

Yes. Median result per workload, latest release, medium effort. Negative is better.

ModelQuery tuning: reads / timeIndex tuning: reads / timeCode review: reads / time
GPT-5.6 Luna−94.1% / −97.3%−78.8% / −38.6%−69.7% / −33.7%
GPT-6 Sol−95.0% / −99.0%−78.7% / −44.3%−25.9% / −21.4%

Zero regressions and zero wrong results across all 35 runs on the latest release. Query tuning is the spectacular case: with the execution plan in hand, both models took a slow report query to nearly nothing. The OLTP row is the one to weigh commercially — an already well-indexed transactional workload got about 79% cheaper and roughly 40% faster.

One behavior DBAs will appreciate: the latest release won’t recommend dropping an “unused” index until the server has been up for at least a week. Our benchmark VM had been up four days, and the index-tuning reports held back their drop suggestions and said why. Usage counters reset on restart, so four days of zero reads doesn’t prove an index is dead.

Question 2 — Does the model matter, and which is best?

Head to head on the latest release:


GPT-5.6 Luna
GPT-6 SolVerdict
Index tuning: time cut−38.6%−44.3%Tie — 5.7 points apart, noise floor 12.4%
Query tuning: time cut−97.3%−99.0%Tie — both near 100%
Code review: time cut−33.7%−21.4%Luna â€” 12.3 points apart, noise floor 10.3%
Code review: reads cut−69.7%−25.9%Luna
Time to answer44–54 s50–59 sTie
Full-strength index advice3 of 5 runs4 of 5 runsSol, slightly more consistent
Wrong results0 of 150 of 15Tie
Bar graph comparing the performance of gpt-5. 6 luna and gpt-6 sol on index tuning, query tuning, and code review, highlighting logical reads and total workload time.
Does ai actually tune sql server? What our ai sql tuning benchmark found 5

Code review is where the models genuinely differ. GPT-5.6 Luna consistently found the fixes that cut the most reads, while GPT-6 Sol at medium often settled for smaller ones. Five runs each isn’t enough for a strong statistical claim on its own (p = 0.15). But the gap is larger than the noise floor, and a separate run on a pre-release build showed the same gap (−39.7% vs. −24.0%). Two separate runs pointing the same way is what we trust.

Index tuning is where GPT-6 Sol earns its keep. The difference in median benefit sits inside the noise floor, so we can’t call it a win. The spread tells a different story, though: Sol’s weakest run still cut total workload time by 35%, against 22% for Luna’s. Luna occasionally proposes just the single most obvious index and stops; Sol did that less often. Sol also added more index storage to get there — 37 MB against 21 MB.

Our recommendation: GPT-5.6 Luna for most work, and certainly for code review. It’s the better code reviewer, it ties on everything else, and it costs a fraction as much. Reach for GPT-6 Sol when index tuning consistency matters more than cost.

Question 3 — How much does reasoning effort matter?

Effort clearly matters on code review. Median cut in total workload time, with median time to answer:

ModelLowMediumHigh
GPT-5.6 Luna−11.6% · 31 s−39.7% · 44 s−39.8% · 120 s
GPT-6 Sol−13.9%² · 32 s−24.0% · 51 s−30.6% · 97 s

² Two of GPT-6 Sol’s five low-effort runs returned wrong results — a rewritten procedure that SQL Server refused to run. Those runs are excluded from the median.

Bar graph comparing the performance of gpt-5. 6 luna and gpt-6 sol in code reviews, highlighting median time to answer for low, medium, and high effort levels.
Does ai actually tune sql server? What our ai sql tuning benchmark found 6

Low is too low for code review. GPT-5.6 Luna split in two at low effort: two runs found the full fix (about −40%), and three settled for a minor one (−2% to −12%). At medium, all five found it. GPT-6 Sol’s two wrong results at low were its only wrong results anywhere in the benchmark.

High buys little over medium. For GPT-5.6 Luna, high and medium were identical (−39.8% vs. −39.7%), but high took almost three times as long. For GPT-6 Sol, high did go after the larger read savings more often, though the gain in total time stayed inside the noise floor. Query tuning told the same story: GPT-6 Sol scored −99.0% at medium and −98.7% at high (p = 0.90), with high taking 89 seconds against 59.

A note on provenance: this effort comparison ran on a pre-release build, one prompt update before the latest release. Medium was then re-run on the latest release and landed within the noise floor of these numbers for both models.

That’s why medium is the default reasoning effort in AI SQL Tuner Studio, including in the Free edition.

What we measured

Every run recorded all of the following, and the report shows them side by side rather than blended:

MetricWhy it’s there
Logical reads, read queriesHeadline. Deterministic, tightest noise floor
CPU time, read queriesSecondary — real, but noisier
Total workload timeThe net result: weighted elapsed time of every read and write, before vs. after
Write cost (reads, CPU, elapsed)The bill for new indexes. Reported beside the benefit, never summed into it
Net index storage added or reclaimedA 90% read reduction that costs 2 GB is a different recommendation from one that costs nothing
Scripts offered / applied / failed / refusedMore recommendations is not better
Time to get the adviceWhat the person at the keyboard actually waits for
Result hash matchPass/fail gate. Changed results void the run
Tokens in and outA check that no response was cut short

Limitations — read these before quoting the numbers

This is one experiment, not a law of nature. Specifically:

  • Two models through the full benchmark. The wider field of eight was screened on an earlier release, and the eliminated models were not re-run on the latest one.
  • Medium only, for index tuning. Index tuning was measured at medium effort only; code review’s effort comparison ran on a pre-release build, as noted above.
  • One environment. A dedicated Azure VM (4 vCPU, 32 GB RAM, premium SSD) running SQL Server 2022 Enterprise, one instance, nothing else on the box. All figures are relative improvements; we publish no absolute timings.
  • One database size. About 1.4 GB, comfortably resident in memory. I/O-bound workloads and much larger tables may behave differently.
  • One schema. A Q&A-site data model and its generated data.
  • A short history. Query Store holds 40 passes of workload history — tens of thousands of executions, but compressed into under a minute rather than the weeks a production server accumulates.
  • A specific read/write mix. Roughly 11% writes by count on index tuning and code review; query tuning is read-only.
  • Statistics were frozen and auto-statistics disabled. Necessary for clean measurement, but not how production runs.
  • Correctness was checked on the workload’s own parameter values. A rewrite that’s wrong only for inputs we didn’t test would pass.
  • Advice we can’t measure isn’t scored. SQL injection risk, SELECT *, naming and maintainability may all be flagged correctly and get no credit here.
  • Five replicates per model per test. That supports descriptive statements, not strong claims of significance. Where we make a statistical claim, we name the p-value.
  • This measures the tool driven by each model, not the models in general. Different prompts would produce different results, and every model is a moving target.
  • This is not a database engine benchmark. We held SQL Server constant and varied the AI.

What we did with it

The results shaped the product directly:

  • The lineup is GPT-5.6 Luna and GPT-6 Sol. Luna is the default in every edition; Sol is the premium option.
  • Medium is the default reasoning effort, and it’s available in the Free edition.
  • Query Tuner sends the execution plan, and index tuning waits for a week of uptime before suggesting an index drop.

If you want to see the kind of output being measured here, the sample reports show the real thing — the same format the harness parses, applies and grades. And if you’d rather generate your own and check our work against your own workload, pick an edition and run it. Read-only permissions on the SQL Server side are still all that’s required.

Questions about the methodology, or a workload you think would break it? We’d genuinely like to hear about it: support@aisqltuner.com or Contact AI SQL Tuner — AI SQL Tuner Studio Support.

GPT is a trademark of OpenAI. Claude and Claude Opus are trademarks of Anthropic, PBC. AI SQL Tuner Studio is a product of AI SQL Tuner LLC.

AI SQL Tuner

Thank You, we'll be in touch soon.
AI SQL Tuner Studio - SQL Server tuning for DBAs and devs, powered by AI. | Product Hunt

© 2026 AI SQL Tuner LLC · AI-Powered SQL Server Optimization. All rights reserved.