ModelVerdict

Comparison

GPT-6 Sol vs Qwen3.8 Flash

Same documents, same prompt: where each model is right, what it really costs and how fast it answers.

At a glance

Task · documentsEnglish (US)CzechGermanSpanishFrench
Invoice extraction99.9% / 82.4%99.9% / 91.8%100.0% / 93.5%99.2% / 87.9%99.9% / 90.6%
Email triage99.0% / 97.5%99.0% / 98.5%98.2% / 97.9%——
Personal data detection99.7% / 99.2%99.7% / 97.0%99.7% / 99.6%——
Contract clauses100.0% / 95.1%100.0% / 99.9%100.0% / 98.0%——

Correct fields: GPT-6 Sol / Qwen3.8 Flash; the better one in bold. Click a cell to see the details.

Invoice extraction · English (US) (150 documents)

GPT-6 Sol is more accurate by 17.6 pp. Qwen3.8 Flash is 8.4× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.9%82.4%
Error-free documents99%57%
Per 1,000 documents$9.65$1.15
To fix / 1,0007427
Speed (median)4.1 sec17.4 sec
Valid answers100.0%90.7%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Supplier100.0%85.3%
Supplier ID (EIN, IČO…)100.0%88.7%
Invoice number100.0%85.3%
Payment reference100.0%79.3%
Invoice date100.0%85.3%
Due date100.0%85.3%
Net by tax rate100.0%70.0%
Tax by rate100.0%76.0%
Total100.0%85.3%
Currency100.0%84.0%
Bank account99.3%81.3%

Documents where they differ

Invoice extraction · Czech (150 documents)

GPT-6 Sol is more accurate by 8.1 pp. Qwen3.8 Flash is 9.9× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.9%91.8%
Error-free documents99%77%
Per 1,000 documents$10.12$1.02
To fix / 1,00013233
Speed (median)3.9 sec17.0 sec
Valid answers100.0%96.7%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Supplier99.3%93.3%
Supplier ID (EIN, IČO…)100.0%94.0%
VAT ID100.0%94.0%
Invoice number100.0%94.7%
Payment reference99.3%94.0%
Invoice date100.0%94.7%
Tax point100.0%94.7%
Due date100.0%94.7%
Net by tax rate100.0%80.0%
Tax by rate100.0%80.7%
Total100.0%94.0%
Currency100.0%94.0%
Bank account100.0%90.7%

Documents where they differ

Invoice extraction · German (150 documents)

GPT-6 Sol is more accurate by 6.5 pp. Qwen3.8 Flash is 9.5× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields100.0%93.5%
Error-free documents99%87%
Per 1,000 documents$10.65$1.12
To fix / 1,0007127
Speed (median)4.9 sec19.0 sec
Valid answers100.0%94.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Supplier100.0%94.0%
Supplier ID (EIN, IČO…)100.0%94.0%
VAT ID100.0%94.0%
Invoice number100.0%92.7%
Payment reference100.0%88.7%
Invoice date100.0%94.0%
Tax point100.0%94.0%
Due date100.0%94.0%
Net by tax rate100.0%94.0%
Tax by rate100.0%94.0%
Total100.0%94.0%
Currency100.0%94.0%
Bank account99.3%94.0%

Documents where they differ

Invoice extraction · Spanish (100 documents)

GPT-6 Sol is more accurate by 11.2 pp. Qwen3.8 Flash is 7.7× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.2%87.9%
Error-free documents89%74%
Per 1,000 documents$10.76$1.40
To fix / 1,000110260
Speed (median)5.7 sec22.3 sec
Valid answers100.0%90.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Supplier98.0%86.0%
Supplier ID (EIN, IČO…)100.0%89.0%
VAT ID100.0%86.0%
Invoice number100.0%89.0%
Payment reference100.0%85.0%
Invoice date100.0%89.0%
Tax point100.0%87.0%
Due date100.0%89.0%
Net by tax rate92.0%86.0%
Tax by rate100.0%90.0%
Total100.0%89.0%
Currency100.0%89.0%
Bank account99.0%89.0%

Documents where they differ

Invoice extraction · French (100 documents)

GPT-6 Sol is more accurate by 9.3 pp. Qwen3.8 Flash is 9× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.9%90.6%
Error-free documents99%82%
Per 1,000 documents$11.05$1.23
To fix / 1,00010180
Speed (median)5.9 sec19.8 sec
Valid answers100.0%95.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Supplier100.0%91.0%
Supplier ID (EIN, IČO…)100.0%91.0%
VAT ID100.0%91.0%
Invoice number100.0%90.0%
Payment reference99.0%86.0%
Invoice date100.0%91.0%
Tax point100.0%92.0%
Due date100.0%89.0%
Net by tax rate100.0%91.0%
Tax by rate100.0%94.0%
Total100.0%91.0%
Currency100.0%91.0%
Bank account100.0%90.0%

Documents where they differ

Email triage · English (US) (150 documents)

GPT-6 Sol is more accurate by 1.6 pp. Qwen3.8 Flash is 8× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.0%97.5%
Error-free documents92%79%
Per 1,000 documents$2.44$0.30
To fix / 1,00080213
Speed (median)2.8 sec8.9 sec
Valid answers100.0%100.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Category100.0%100.0%
Priority98.0%97.3%
Sentiment93.3%80.7%
Customer name100.0%100.0%
Order number100.0%100.0%
Invoice number100.0%100.0%
Amount100.0%100.0%
Currency100.0%100.0%
Deadline100.0%99.3%

Documents where they differ

Email triage · Czech (150 documents)

GPT-6 Sol is more accurate by 0.4 pp. Qwen3.8 Flash is 7.2× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.0%98.5%
Error-free documents91%87%
Per 1,000 documents$3.19$0.45
To fix / 1,00093133
Speed (median)2.8 sec13.0 sec
Valid answers100.0%100.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Category100.0%100.0%
Priority98.0%99.3%
Sentiment92.7%88.7%
Customer name100.0%99.3%
Order number100.0%99.3%
Invoice number100.0%100.0%
Amount100.0%100.0%
Currency100.0%100.0%
Deadline100.0%100.0%

Documents where they differ

Email triage · German (150 documents)

GPT-6 Sol is more accurate by 0.4 pp. Qwen3.8 Flash is 7× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields98.2%97.9%
Error-free documents84%81%
Per 1,000 documents$2.73$0.39
To fix / 1,000160187
Speed (median)2.9 sec12.1 sec
Valid answers100.0%100.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Category100.0%100.0%
Priority97.3%96.0%
Sentiment86.7%84.7%
Customer name100.0%100.0%
Order number100.0%100.0%
Invoice number100.0%100.0%
Amount100.0%100.0%
Currency100.0%100.0%
Deadline100.0%100.0%

Documents where they differ

Personal data detection · English (US) (150 documents)

GPT-6 Sol is more accurate by 0.5 pp. Qwen3.8 Flash is 7.5× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.7%99.2%
Error-free documents98%99%
Per 1,000 documents$2.09$0.28
To fix / 1,0002013
Speed (median)2.6 sec7.5 sec
Valid answers100.0%100.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Names100.0%99.3%
Email addresses100.0%99.3%
Phone numbers100.0%99.3%
Addresses98.0%98.7%
Dates of birth100.0%99.3%
National IDs100.0%99.3%
Bank accounts100.0%99.3%

Documents where they differ

Both models got the same number of fields right on every example document.

Personal data detection · Czech (150 documents)

GPT-6 Sol is more accurate by 2.8 pp. Qwen3.8 Flash is 5.5× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.7%97.0%
Error-free documents98%89%
Per 1,000 documents$2.22$0.41
To fix / 1,00020113
Speed (median)2.5 sec9.6 sec
Valid answers100.0%99.3%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Names100.0%90.7%
Email addresses100.0%98.0%
Phone numbers100.0%97.3%
Addresses98.0%96.7%
Dates of birth100.0%98.0%
National IDs100.0%99.3%
Bank accounts100.0%98.7%

Documents where they differ

Personal data detection · German (150 documents)

GPT-6 Sol is more accurate by 0.1 pp. Qwen3.8 Flash is 5.9× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields99.7%99.6%
Error-free documents98%99%
Per 1,000 documents$2.13$0.36
To fix / 1,000207
Speed (median)2.6 sec12.2 sec
Valid answers100.0%100.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Names100.0%99.3%
Email addresses100.0%99.3%
Phone numbers100.0%99.3%
Addresses98.0%99.3%
Dates of birth100.0%100.0%
National IDs100.0%100.0%
Bank accounts100.0%100.0%

Documents where they differ

Both models got the same number of fields right on every example document.

Contract clauses · English (US) (150 documents)

GPT-6 Sol is more accurate by 4.9 pp. Qwen3.8 Flash is 5.4× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields100.0%95.1%
Error-free documents100%69%
Per 1,000 documents$3.33$0.62
To fix / 1,0000307
Speed (median)2.5 sec14.1 sec
Valid answers100.0%100.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Provider100.0%69.3%
Provider ID100.0%100.0%
Client100.0%69.3%
Client ID100.0%100.0%
Contract type100.0%100.0%
Price (net)100.0%100.0%
Currency100.0%100.0%
Billing period100.0%100.0%
Effective date100.0%99.3%
End date100.0%100.0%
Notice period (months)100.0%100.0%
Late-payment penalty (% per day)100.0%99.3%
Governing law100.0%99.3%

Documents where they differ

Contract clauses · Czech (150 documents)

GPT-6 Sol is more accurate by 0.1 pp. Qwen3.8 Flash is 7.3× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields100.0%99.9%
Error-free documents100%99%
Per 1,000 documents$4.62$0.63
To fix / 1,000013
Speed (median)2.0 sec13.6 sec
Valid answers100.0%100.0%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Provider100.0%100.0%
Provider ID100.0%100.0%
Client100.0%100.0%
Client ID100.0%100.0%
Contract type100.0%98.7%
Price (net)100.0%100.0%
Currency100.0%100.0%
Billing period100.0%100.0%
Effective date100.0%100.0%
End date100.0%100.0%
Notice period (months)100.0%100.0%
Late-payment penalty (% per day)100.0%100.0%
Governing law100.0%100.0%

Documents where they differ

Contract clauses · German (150 documents)

GPT-6 Sol is more accurate by 2 pp. Qwen3.8 Flash is 5.6× cheaper.

MetricGPT-6 SolQwen3.8 Flash
Correct fields100.0%98.0%
Error-free documents100%87%
Per 1,000 documents$4.36$0.78
To fix / 1,0000127
Speed (median)2.9 sec19.7 sec
Valid answers100.0%99.3%

Accuracy per field

FieldGPT-6 SolQwen3.8 Flash
Provider100.0%99.3%
Provider ID100.0%94.7%
Client100.0%99.3%
Client ID100.0%94.7%
Contract type100.0%99.3%
Price (net)100.0%99.3%
Currency100.0%99.3%
Billing period100.0%99.3%
Effective date100.0%99.3%
End date100.0%99.3%
Notice period (months)100.0%99.3%
Late-payment penalty (% per day)100.0%99.3%
Governing law100.0%91.3%

Documents where they differ