> The problem is that the differences between flagship and local models are compounding heavily
This depends a lot on how you work, and how much of the architectural thinking you do yourself.
People seem to lose sight of the fact that a flash model today is as powerful as a frontier model from a year ago. If you were happy with GPT 4.x, you should be ecstatic that equivalent power is now basically free...
I find that with a lot of the cheaper models, I end up spending a lot more time correcting the easy stuff.
If I am 100% spot on on the architectural stuff, I have anecdata that some of the frontier models might actually be cheaper than "cheaper" alternatives once you look at what it takes to get to good output, since they require less correction.
But that is on pure token costs. When you value the human overseer's time, there is just no competition. A model that is 10x more expensive that requires 10% less oversight is just a plain win.
This depends a lot on how you work, and how much of the architectural thinking you do yourself.
People seem to lose sight of the fact that a flash model today is as powerful as a frontier model from a year ago. If you were happy with GPT 4.x, you should be ecstatic that equivalent power is now basically free...