Ashman Das

CASCaDER

Paused since 20 September 2026. I plan to pick it up again around February 2027.

Most language models spend the same amount of work on every word they read, whether it’s “the” or the hard part of a maths problem. CASCaDER asks whether a model can learn to spend less on easy words and more on hard ones, at the same total cost.

Answer loss falls from 0.0394 at 1 loop to 0.0167 at 4 loops, then stays flat at 8 and 16 loops. 0.020.030.04 124816 1 loop: loss 0.03942 loops: loss 0.02194 loops: loss 0.01678 loops: loss 0.016916 loops: loss 0.017 best
Loss on the easy task by number of loops, lower is better. After 4 loops nothing changes.

The three parts

The part that decides the budget is designed to watch the last few tokens, not just the current one. So a token’s budget depends on what the text around it needed. Most work on dynamic compute decides from each token alone.

What I measured

All runs were on one laptop RTX 4060 with 8 GB of memory. The model has about 108 million parameters in total, of which only 4 to 18 million are active for any one token.

Looping more doesn’t help after four

I trained a model that could loop up to 16 times, then tested it at different loop counts. The best result came at 4 loops on both tasks. Going to 8 or 16 changed nothing. I wrote the pass-or-fail rule into the code before running it, and it failed. That saved a planned 40 GPU-hour follow-up, which a 45-second test showed wasn’t needed.

Width and depth tie at the same compute

Spending the same compute on more experts (width) or more loops (depth) gave the same model. Across three seeds each, the loss agreed to 0.0001. But the looped version cost 2.16 times more at inference, because loops run one after another.

Whether width and depth substitute for each other

A small grid of width × depth settings suggested that once the model loops, extra width stops helping. That grid only finished one seed before I stopped it, so I don’t count it as a result.

How I keep myself honest

Negative results stay in the record under their real name. Three sets of results I later found were broken, including one where future tokens leaked into training, live in a folder called results/invalid instead of being deleted. In one pilot I found that my own headline number could be predicted by a simple formula (0.286 predicted, 0.259 measured), which meant the model was taking a shortcut. I wrote that down too.

Why it’s paused

The last stretch ran for two weeks straight on my only computer. Each failed run cost days, so I was learning very little per week. I’ll restart when a failed experiment costs hours instead of days.