🔬 MuCoCo Benchmark Results

← Back to Home

Model Performance Across Mutation Types

Task Performance Across Mutation Types

Dataset Performance Across Mutation Types

Effectiveness of MuCoCo Across All Models

Mutation Type Qwen2.5 Gemma-3 DeepSeek-V3.2 LLaMA-3.1 GPT-5 GPT-4o Codestral All Models
Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc.
Logical 7.60 76.23 12.24 66.13 15.72 63.00 17.21 49.88 1.52 97.52 17.99 64.25 12.67 66.15 11.38 (3829/33646) 69.01 (310553/450000)
Syntactic 4.19 77.91 7.92 68.78 12.95 65.80 11.43 55.70 1.10 98.46 15.81 65.75 9.09 68.71 8.38 (14608/174133) 71.58 (1166123/23909)
Lexical 13.19 67.63 17.93 60.66 12.92 69.24 22.17 48.41 2.78 94.16 17.05 68.90 10.69 70.96 13.07 (7607/58213) 68.49 (534425/794644)
All (Avg.) 9.92 71.86 14.47 63.60 13.76 66.80 18.89 50.00 2.11 95.87 17.14 66.99 11.04 69.14 11.79 (12896/109272) 69.13 (1020992/1476773)

Effectiveness of MuCoCo across all tested tasks

Mutation Type MCQ Input Pred. Output Pred. Code Gen. All Tasks
Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc.
Logical 45.56 46.48 8.78 75.90 10.77 64.16 11.38 (3829/33646) 69.01 (310553/450000)
Syntactic 19.82 67.80 7.02 80.53 9.08 62.89 8.38 (1460/17413) 71.58 (16612/23209)
Lexical 10.36 78.36 7.88 80.59 12.38 61.84 26.22 57.68 13.07 (7607/58213) 68.49 (54425/79464)
All (Avg.) 25.60 64.04 8.03 78.97 11.23 62.82 26.22 57.68 11.79 (12896/109272) 69.13 (102092/147673)

Effectiveness of MuCoCo across coding datasets

Mutation Type CodeMMLU HumanEval CruxEval BigCodeBench All Datasets
Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc. Inc. Acc.
Logical 45.56 46.48 8.41 73.26 14.06 60.57 11.38 (3829/33646) 69.01 (310553/450000)
Syntactic 19.82 67.80 6.70 74.94 10.00 66.74 8.38 (1460/17413) 71.58 (16612/23209)
Lexical 10.36 78.36 9.00 74.55 11.82 66.41 26.98 56.97 13.07 (7607/58213) 68.49 (54425/79464)
All 25.60 64.04 8.35 74.16 11.97 65.11 26.98 56.97 11.79 (12896/109272) 69.13 (102092/147673)

Effectiveness of MuCoCo versus TURBULENCE using the TURBULENCE dataset

Approach Code Generation Input Prediction Output Prediction All Tasks
#errs #tests ErrRate #errs #tests ErrRate #errs #tests ErrRate #errs #tests ErrRate
TURBULENCE 17604 40567 43.39% 54752 134169 40.81% 63651 225613 28.21% 136007 400349 33.97%
MuCoCo 23933 73279 32.66% 366968 1036344 35.4% 405316 1490118 27.2% 796217 2599741 30.63%
% Impr. 35.95% 80.64% -24.73% 570.24% 672.42% -13.26% 536.78% 560.48% -3.58% 485.42% 549.37% -9.83%

Impact of varying mode confidence on Inconsistency (%) and Accuracy (%)

Confidence Gemma Qwen Llama Aggregated
HumanEval Inc. HumanEval Acc. CruxEval Inc. CruxEval Acc. CodeMMLU Inc. CodeMMLU Acc. HumanEval Inc. HumanEval Acc. CruxEval Inc. CruxEval Acc. CodeMMLU Inc. CodeMMLU Acc. HumanEval Inc. HumanEval Acc. CruxEval Inc. CruxEval Acc. CodeMMLU Inc. CodeMMLU Acc. Inc. Acc.
0.50 18.02 82.28 4.86 77.09 51.21 52.52 3.39 88.52 7.30 84.21 29.93 63.10 18.02 47.54 4.86 93.94 51.21 37.04 9.71 75.79
0.60 18.02 82.28 4.86 77.09 50.00 52.52 3.39 88.52 7.30 84.21 29.93 63.10 18.02 47.54 4.86 93.94 50.00 37.74 9.65 75.86
0.70 18.02 82.28 4.86 77.09 36.15 52.52 3.39 88.52 7.30 84.21 29.38 64.03 18.02 47.54 4.86 93.94 36.15 42.83 9.15 76.33
0.80 0.25 84.08 0.24 78.45 8.54 53.05 0.07 95.88 0.68 92.65 24.38 67.95 0.25 48.13 0.24 98.43 8.54 55.50 3.91 82.79
0.90 0.58 85.03 0.00 79.43 0.00 54.50 0.00 98.42 0.15 97.47 18.99 74.34 0.58 55.47 0.00 99.70 0.00 82.43 3.59 86.22
0.95 0.00 86.31 0.00 80.92 0.00 55.33 0.00 99.38 0.28 99.25 14.91 79.75 0.00 66.52 0.00 100.00 0.00 88.24 3.35 86.74
0.99 0.00 89.68 0.00 85.74 0.00 56.54 0.00 100.00 0.00 100.00 4.11 88.75 0.00 100.00 0.00 100.00 0.00 100.00 2.06 88.32