Context : We continuously do runs of a subset of 64 curated SWE-Bench Pro long horizon tasks on various models. Some via their respective API's. Some on reference hardware setups. We do this to approximate 'real world use' of these models and get more insights on how these use the underlying compute (we advise HW builders)
In our runs we create multiple separate sandboxed agents that use a given model/harness combo (in this case just the default harnesses for both) and feed them the benchmark tasks, which we compare against the golden resolution. (There's also a time out, just to make sure we don't blow through our entire budget by accident). You pay for the tokens no matter if the task gets solved or not, so both bills cover all 64 attempts, also the failed ones (30 for Astra, 13 for Fable). The €68 is just the difference between the two full runs, not the price of those 17 tasks. That's basically why we look at cost per solved task, €1.6 vs €2.4. But we see how the wording of that sentence wasn't optimal. We'll fix that bullet (thx)
Aside from it giving us a good view on evolving token needs of different model generations (and being able to compare with open weights models), the variations in "intelligence per dollar" per generation is also pretty interesting. Hence posts like these, just to share the data, which is hopefully useful to others too. And : always interested in seeing different results.
> exist. A new study found that everyday people generally do not find these targeted words offensive, and reading them in a story does not negatively affect psychological well-being.
There’s only a certain kind of person who would find “master bedroom” offensive or harmful.
Some of them may be black, admittedly, but mostly it's white people being offended on behalf of black people. And deaf people. And non-neurotypical people. And visually impaired people. And ... It's rather a labour-intensive hobby to engage in.
> It's very frustrating talking about this subject with Americans online[1]. In general they seem to have blinders on regarding the magnitude of the consequences of their democratic choice in leadership.
If polls are to be believed, the majority of Americans do not, in fact, fit that description.
> This seems like a pretty catastrophic failure on Amazon’s part. It seems pretty incompetent actually. You should have multiple copies of your customer’s data across multiple locations
How much of it is due to Amazon VS due to regulatory reasons (have to host data in particular region) VS customer self-serve have to choose redundancy across regions?
My quick read through the article didn’t surface an answer to this, but it is very possible I missed it.
So Astra billed for tasks it couldn’t solve?
reply