Hacker Newsnew | past | comments | ask | show | jobs | submit | andsoitis's commentslogin

> Fable solved 17 more tasks than Astra, but also billed €68 more for them.

So Astra billed for tasks it couldn’t solve?


Correct. But that's by design here.

Context : We continuously do runs of a subset of 64 curated SWE-Bench Pro long horizon tasks on various models. Some via their respective API's. Some on reference hardware setups. We do this to approximate 'real world use' of these models and get more insights on how these use the underlying compute (we advise HW builders)

In our runs we create multiple separate sandboxed agents that use a given model/harness combo (in this case just the default harnesses for both) and feed them the benchmark tasks, which we compare against the golden resolution. (There's also a time out, just to make sure we don't blow through our entire budget by accident). You pay for the tokens no matter if the task gets solved or not, so both bills cover all 64 attempts, also the failed ones (30 for Astra, 13 for Fable). The €68 is just the difference between the two full runs, not the price of those 17 tasks. That's basically why we look at cost per solved task, €1.6 vs €2.4. But we see how the wording of that sentence wasn't optimal. We'll fix that bullet (thx)

Aside from it giving us a good view on evolving token needs of different model generations (and being able to compare with open weights models), the variations in "intelligence per dollar" per generation is also pretty interesting. Hence posts like these, just to share the data, which is hopefully useful to others too. And : always interested in seeing different results.


I’d rather pay more and get things solved.

Penny wise and pound foolish comes to mind.


Anything other than cow's milk is a poor imitation.

This makes no sense.

Looks like the photo was taken outside 630 George Street, Sydney NSW 2000, looking at the building's George Street frontage.

Yes. Don't succumb to featuritis. Destroy the barnacles.

> exist. A new study found that everyday people generally do not find these targeted words offensive, and reading them in a story does not negatively affect psychological well-being.

There’s only a certain kind of person who would find “master bedroom” offensive or harmful.


Yep, they are black. You can just say it vs hinting at it;)

Not according to my observations...

No, they're SJWs.

Some of them may be black, admittedly, but mostly it's white people being offended on behalf of black people. And deaf people. And non-neurotypical people. And visually impaired people. And ... It's rather a labour-intensive hobby to engage in.


In my experience that’s not it.

> It's very frustrating talking about this subject with Americans online[1]. In general they seem to have blinders on regarding the magnitude of the consequences of their democratic choice in leadership.

If polls are to be believed, the majority of Americans do not, in fact, fit that description.


If the elections are to believed, a majority of Americans don't care enough to vote against Trump.

> It's a weird two-front war

Who are the adversaries?


> This seems like a pretty catastrophic failure on Amazon’s part. It seems pretty incompetent actually. You should have multiple copies of your customer’s data across multiple locations

How much of it is due to Amazon VS due to regulatory reasons (have to host data in particular region) VS customer self-serve have to choose redundancy across regions?

My quick read through the article didn’t surface an answer to this, but it is very possible I missed it.


When someone says "Trust Me"...

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: