Pricing products with a 3B model on one consumer GPU

Llama-3.2-3B, quantized to 4 bits, with a LoRA adapter trained on a single 16 GB desktop GPU. The reference recipe for this configuration asks for a paid A100 with high memory.

Mean error after fine-tuning
$38.61
Before
$144.47
Improvement
3.7×
Source on GitHub Adapter on the Hub 250 held-out test rows 389 M trainable params RTX 5070 Ti (16 GB)
This page is not a live demo. Hugging Face gates Gradio Spaces behind a PRO subscription, and no free host runs a 3 B model on a GPU. The interactive app is in app.py in the repository and runs locally. Every number below was measured on the machine described above, not quoted from anywhere.

Fine-tuned against the untuned base

Both models scored by the same harness on the same 250 rows. The base-model column is the honest control: it separates what the fine-tune taught from what Llama already knew about prices.

$38.61
Mean error, fine-tuned
$144.47
Mean error, base
73.1%
R², fine-tuned
-406.9%
R², base

The striking figure is the base model's R² of -406.9%. Negative R² means it scores worse than a constant predictor that guesses the average price every time. The next section shows why.

Every model scored, same rows, same harness

Hosted frontier baselines are not included. Google's free tier now caps Gemini at 10–20 requests per day per model, which is far short of the 250 needed to compare at the same sample size, and quoting someone else's benchmark as if it were measured here would defeat the point. pricer/baselines.py runs them on a paid key.

ModelMean errorMedian R²Within 20%Runs on
Fine-tuned Llama-3.2-3B (QLoRA) ← $38.61 $17.99 73.1% 44.8% local, 4-bit, fine-tuned
Base Llama-3.2-3B (4-bit) $144.47 $40.80 -406.9% 12.4% local, 4-bit, untuned

The base model recites prices; it does not estimate them

Across 250 products the untuned model produced only 27 distinct answers, and its five most frequent replies account for 76% of everything it said — $99.99 ×78, $9.99 ×47, $12.99 ×26. It answers with common internet price strings largely independent of the product in front of it. Fine-tuning roughly doubles that variety to 53 distinct answers with the top five covering 55%, and, more importantly, ties the answer to the item.

The same weights, with and without the adapter

A LoRA adapter is a small set of matrices layered over the frozen base model, so it can be switched off in place — the same loaded weights answer twice. These are real answers on recognisable products, which is the nearest thing to a live demo this page can offer:

ProductCategory Adapter onAdapter off
Apple iPhone 15 Pro 128GB Smartphone Electronics $900.00 $1,000.00
Sony WH-1000XM5 Wireless Noise Cancelling Headphones Electronics $298.00 $299.99
LEGO Star Wars Millennium Falcon Building Set Toys & Games $200.00 $99.99
Bosch ICON Beam Windshield Wiper Blades, Pair Automotive Accessories $20.00 $9.99

Read the right-hand column on its own. Every answer is a stock price ending in .99 or a round thousand — the same handful of strings the untuned model reaches for whatever it is shown. The left column is the same network, same quantization, same prompt, with 389 M adapter parameters layered on top.

Error by price band

A single average can hide a model that is only good on cheap items, so the same errors split by true price:

True pricen BaseFine-tunedBetter by
$0–$25 41 $10.46 $6.78 1.5×
$25–$100 102 $45.52 $17.15 2.7×
$100–$300 84 $223.19 $47.08 4.7×
$300+ 23 $534.69 $159.60 3.4×

The gap widens as items get more expensive, which is what you would expect if the base model is anchoring on cheap round numbers.

Actual predictions

Sampled evenly across the price range rather than hand-picked — sorted by true price, then stepped through at fixed intervals, so the weak cases are included:

ProductActual Fine-tunedBase
Retevis RT68 Soft Earhook Walkie Talkie ... $6.99 $20.00 $9.99
Berta 1-1/4" Overlay 90° Soft Close Cabi... $24.66 $30.00 $12.99
ReplacementBrand 3-Pack GE RPWF Compatib... $36.99 $30.00 $9.99
Forged Carbon Fiber Case for iPhone 14 P... $59.99 $40.00 $99.99
Sapphire Radeon HD 5450 1 GB PCI‑Express... $99.99 $129.00 $99.99
Vinyl Roll‑Up Tonneau Cover Kit for Toyo... $148.88 $159.00 $199.99
MGP Caliper Covers – Set of 4 Red Powder... $249.00 $249.00 $24.99
Dell XPS 13 13.3" Full‑HD Laptop $839.99 $500.00 $1,000.00

What was hard

Windows killed the run at step 607. Training died inside loss.backward() with CUDA error: unknown error. The cause was TDR — Windows resets the display driver if a GPU operation appears hung for longer than TdrDelay, which defaults to two seconds. A rank-256 backward pass on a card that also drives monitors exceeds that. No ECC errors, no retired pages: the hardware was fine, the default was simply wrong for compute.

Then it got 9× slower with no error at all. Training needs about 15 GB of 16.3 GB. When a few hundred megabytes of desktop applications tipped the allocation over, the driver silently paged to system RAM over PCIe instead of raising a CUDA OOM. The tell is counter-intuitive: GPU utilization pinned at 100% while power draw sat at 114 W of 300 W. High utilization with low power means the GPU is waiting on the bus, not computing — utilization alone never reveals it, only step time does. The fix was a smaller micro-batch at an identical effective batch of 256, trading 55% throughput for a run that finished unattended.

Made by Yug