CASCaDER
Paused since 20 September 2026. I plan to pick it up again around February 2027.
Most language models spend the same amount of work on every word they read, whether it’s “the” or the hard part of a maths problem. CASCaDER asks whether a model can learn to spend less on easy words and more on hard ones, at the same total cost.
The three parts
- DAPaRo changes how many experts work on each token. A mixture-of-experts model has many small sub-networks. Normally a fixed number of them handle every token. DAPaRo lets that number vary. This is the main question.
- RDT loops the same block of the model several times before it answers, so it can “think longer” without adding weights. I didn’t invent this. I adopted it because my laptop GPU couldn’t train a model big enough to test DAPaRo any other way.
- COBRA would set one budget for a whole task before the model starts answering. It isn’t built yet.
The part that decides the budget is designed to watch the last few tokens, not just the current one. So a token’s budget depends on what the text around it needed. Most work on dynamic compute decides from each token alone.
What I measured
All runs were on one laptop RTX 4060 with 8 GB of memory. The model has about 108 million parameters in total, of which only 4 to 18 million are active for any one token.
Looping more doesn’t help after four
I trained a model that could loop up to 16 times, then tested it at different loop counts. The best result came at 4 loops on both tasks. Going to 8 or 16 changed nothing. I wrote the pass-or-fail rule into the code before running it, and it failed. That saved a planned 40 GPU-hour follow-up, which a 45-second test showed wasn’t needed.
Width and depth tie at the same compute
Spending the same compute on more experts (width) or more loops (depth) gave the same model. Across three seeds each, the loss agreed to 0.0001. But the looped version cost 2.16 times more at inference, because loops run one after another.
Whether width and depth substitute for each other
A small grid of width × depth settings suggested that once the model loops, extra width stops helping. That grid only finished one seed before I stopped it, so I don’t count it as a result.
How I keep myself honest
Negative results stay in the record under their real name. Three sets of results I later found were broken, including one where future tokens leaked into training, live in a folder called results/invalid instead of being deleted. In one pilot I found that my own headline number could be predicted by a simple formula (0.286 predicted, 0.259 measured), which meant the model was taking a shortcut. I wrote that down too.
Why it’s paused
The last stretch ran for two weeks straight on my only computer. Each failed run cost days, so I was learning very little per week. I’ll restart when a failed experiment costs hours instead of days.