Llama-3.2-3B, quantized to 4 bits, with a LoRA adapter trained on a single 16 GB desktop GPU. The reference recipe for this configuration asks for a paid A100 with high memory.
app.py in the repository and runs
locally. Every number below was measured on the machine described above, not
quoted from anywhere.
Both models scored by the same harness on the same 250 rows. The base-model column is the honest control: it separates what the fine-tune taught from what Llama already knew about prices.
The striking figure is the base model's R² of -406.9%. Negative R² means it scores worse than a constant predictor that guesses the average price every time. The next section shows why.
Hosted frontier baselines are not included. Google's free tier now caps Gemini at 10–20 requests per day per model, which is far short of the 250 needed to compare at the same sample size, and quoting someone else's benchmark as if it were measured here would defeat the point. pricer/baselines.py runs them on a paid key.
| Model | Mean error | Median | R² | Within 20% | Runs on |
|---|---|---|---|---|---|
| Fine-tuned Llama-3.2-3B (QLoRA) ← | $38.61 | $17.99 | 73.1% | 44.8% | local, 4-bit, fine-tuned |
| Base Llama-3.2-3B (4-bit) | $144.47 | $40.80 | -406.9% | 12.4% | local, 4-bit, untuned |
Across 250 products the untuned model produced only
27 distinct answers, and its five most frequent
replies account for 76% of everything it said —
$99.99 ×78, $9.99 ×47, $12.99 ×26. It answers with common internet price strings largely independent of
the product in front of it. Fine-tuning roughly doubles that variety to
53 distinct answers with the top five covering
55%, and, more importantly, ties the answer to the item.
A LoRA adapter is a small set of matrices layered over the frozen base model, so it can be switched off in place — the same loaded weights answer twice. These are real answers on recognisable products, which is the nearest thing to a live demo this page can offer:
| Product | Category | Adapter on | Adapter off |
|---|---|---|---|
| Apple iPhone 15 Pro 128GB Smartphone | Electronics | $900.00 | $1,000.00 |
| Sony WH-1000XM5 Wireless Noise Cancelling Headphones | Electronics | $298.00 | $299.99 |
| LEGO Star Wars Millennium Falcon Building Set | Toys & Games | $200.00 | $99.99 |
| Bosch ICON Beam Windshield Wiper Blades, Pair | Automotive Accessories | $20.00 | $9.99 |
Read the right-hand column on its own. Every answer is a stock price
ending in .99 or a round thousand — the same handful of
strings the untuned model reaches for whatever it is shown. The left column
is the same network, same quantization, same prompt, with 389 M
adapter parameters layered on top.
A single average can hide a model that is only good on cheap items, so the same errors split by true price:
| True price | n | Base | Fine-tuned | Better by |
|---|---|---|---|---|
| $0–$25 | 41 | $10.46 | $6.78 | 1.5× |
| $25–$100 | 102 | $45.52 | $17.15 | 2.7× |
| $100–$300 | 84 | $223.19 | $47.08 | 4.7× |
| $300+ | 23 | $534.69 | $159.60 | 3.4× |
The gap widens as items get more expensive, which is what you would expect if the base model is anchoring on cheap round numbers.
Sampled evenly across the price range rather than hand-picked — sorted by true price, then stepped through at fixed intervals, so the weak cases are included:
| Product | Actual | Fine-tuned | Base |
|---|---|---|---|
| Retevis RT68 Soft Earhook Walkie Talkie ... | $6.99 | $20.00 | $9.99 |
| Berta 1-1/4" Overlay 90° Soft Close Cabi... | $24.66 | $30.00 | $12.99 |
| ReplacementBrand 3-Pack GE RPWF Compatib... | $36.99 | $30.00 | $9.99 |
| Forged Carbon Fiber Case for iPhone 14 P... | $59.99 | $40.00 | $99.99 |
| Sapphire Radeon HD 5450 1 GB PCI‑Express... | $99.99 | $129.00 | $99.99 |
| Vinyl Roll‑Up Tonneau Cover Kit for Toyo... | $148.88 | $159.00 | $199.99 |
| MGP Caliper Covers – Set of 4 Red Powder... | $249.00 | $249.00 | $24.99 |
| Dell XPS 13 13.3" Full‑HD Laptop | $839.99 | $500.00 | $1,000.00 |
Windows killed the run at step 607. Training died inside
loss.backward() with CUDA error: unknown error. The
cause was TDR — Windows resets the display driver if a GPU operation
appears hung for longer than TdrDelay, which defaults to two
seconds. A rank-256 backward pass on a card that also drives monitors exceeds
that. No ECC errors, no retired pages: the hardware was fine, the default was
simply wrong for compute.
Then it got 9× slower with no error at all. Training needs about 15 GB of 16.3 GB. When a few hundred megabytes of desktop applications tipped the allocation over, the driver silently paged to system RAM over PCIe instead of raising a CUDA OOM. The tell is counter-intuitive: GPU utilization pinned at 100% while power draw sat at 114 W of 300 W. High utilization with low power means the GPU is waiting on the bus, not computing — utilization alone never reveals it, only step time does. The fix was a smaller micro-batch at an identical effective batch of 256, trading 55% throughput for a run that finished unattended.
Made by Yug