this post was submitted on 12 Sep 2024
14 points (93.8% liked)

Technology

1184 readers
391 users here now

Which posts fit here?

Anything that is at least tangentially connected to the technology, social media platforms, informational technologies and tech policy.


Rules

1. English onlyTitle and associated content has to be in English.
2. Use original linkPost URL should be the original link to the article (even if paywalled) and archived copies left in the body. It allows avoiding duplicate posts when cross-posting.
3. Respectful communicationAll communication has to be respectful of differing opinions, viewpoints, and experiences.
4. InclusivityEveryone is welcome here regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, education, socio-economic status, nationality, personal appearance, race, caste, color, religion, or sexual identity and orientation.
5. Ad hominem attacksAny kind of personal attacks are expressly forbidden. If you can't argue your position without attacking a person's character, you already lost the argument.
6. Off-topic tangentsStay on topic. Keep it relevant.
7. Instance rules may applyIf something is not covered by community rules, but are against lemmy.zip instance rules, they will be enforced.


Companion communities

[email protected]
[email protected]


Icon attribution | Banner attribution

founded 10 months ago
MODERATORS
 

But it’s not cheap.

top 4 comments
sorted by: hot top controversial new old
[–] [email protected] 2 points 6 days ago

no it doesn't have reasoning abilities. It just replicates you trying to coax it into giving you something decent, hides the process from you, and then charges you for it.

[–] [email protected] 3 points 1 week ago (1 children)

I'm curious how it will do on the private benchmark that ai explained made. I think it was called simple bench?

[–] [email protected] 3 points 1 week ago (1 children)

This is one stat I've found

[–] [email protected] 4 points 1 week ago

https://simple-bench.com/index.html I was referring to this benchmark specifically because the point of it is to benchmark the actual reasoning capabilities of LLMs:

Simple bench is the only reasoning benchmark written in natural language at which English-speaking humans (and yes, even 'smart highschoolers') can score 90%+, while frontier LLMs get less than 50%. It is an encapsulation of the reasoning deficit found in AI like ChatGPT.

These questions are fully private, preventing contamination, and have been vetted by PhDs from multiple domains, as well as the author - Philip, from AI Explained - who first exposed the numerous errors in the MMLU (Aug 2023). This was celebrated by, among others Andrej Karpathy.