sign up
sign up
sign up
sign up
pull down to refresh
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
arxiv.org/abs/2510.04721
210 sats
\
1 comment
\
@jakoyoh629
25 Oct 2025
AI
related
Hallucination Stations On Some Basic Limitations of Transformer-Based LM
arxiv.org/pdf/2507.07505
213 sats
\
0 comments
\
@0xbitcoiner
23 Jan
AI
To Make Language Models Work Better, Researchers Sidestep Language
www.quantamagazine.org/to-make-language-models-work-better-researchers-sidestep-language-20250414/
210 sats
\
0 comments
\
@0xbitcoiner
15 Apr 2025
AI
Large Language Models Pass the Turing Test
arxiv.org/pdf/2503.23674
374 sats
\
11 comments
\
@south_korea_ln
15 Apr 2025
AI
Why language models hallucinate - OpenAI
openai.com/index/why-language-models-hallucinate/
438 sats
\
4 comments
\
@Scoresby
6 Sep 2025
AI
Mathematicians issue a major challenge to AI—show us your work
www.scientificamerican.com/article/mathematicians-launch-first-proof-a-first-of-its-kind-math-exam-for-ai/
1145 sats
\
4 comments
\
@south_korea_ln
14 Feb
AI
science
Researchers discover impressive learning capabilities in long-context LLMs
venturebeat.com/ai/deepmind-researchers-discover-impressive-learning-capabilities-in-long-context-llms/
397 sats
\
0 comments
\
@ch0k1
25 Apr 2024
tech
The human role in AI mathematics: finding the simple proof?
arxiv.org/pdf/2609.02882
1693 sats
\
3 comments
\
@south_korea_ln
3 Sep
AI
math
science
Meet the new biologists treating LLMs like aliens
www.technologyreview.com/2026/01/12/1129782/ai-large-language-models-biology-alien-autopsy/
580 sats
\
1 comment
\
@winteryeti
14 Jan
AI
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection
arxiv.org/abs/2510.04849v1
433 sats
\
2 comments
\
@optimism
19 Oct 2025
AI
Is AI Reasoning Right for the Wrong Reasons?
www.quantamagazine.org/is-ai-reasoning-right-for-the-wrong-reasons-20260731/
253 sats
\
1 comment
\
@0xbitcoiner
2 Aug
AI
science
The ORCA Benchmark Evaluates How Well AIs Deal with Everyday Math
www.omnicalculator.com/reports/omni-research-on-calculation-in-ai-benchmark
260 sats
\
0 comments
\
@0xbitcoiner
27 Feb
AI
An OpenAI model solved a famous math problem that stumped humans for 80 years
arstechnica.com/ai/2026/06/openais-math-breakthrough-played-to-ais-strengths/
338 sats
\
0 comments
\
@0xbitcoiner
1 Jun
AI
math
LLMs and the Specter of the Cognitive Black Hole
www.psychologytoday.com/us/blog/the-digital-self/202403/llms-and-the-specter-of-the-cognitive-black-hole
200 sats
\
0 comments
\
@ch0k1
22 Mar 2024
science
How to turn LLM Pinocchio into a real boy
12.7k sats
\
10 comments
\
@Scoresby
7 Oct 2025
AI
Financial Statement Analysis with Large Language Models
papers.ssrn.com/sol3/papers.cfm?abstract_id=4835311&fbclid=IwY2xjawIJNupleHRuA2FlbQIxMAABHWJxn71ESvZCS0FxEF_31oro1rwtk4rlgOst5Q4A6tuxDhxB9cgZBPizAg_aem_OAMNHiz7Vyv2bb2vt2yM0Q
222 sats
\
2 comments
\
@scatman
31 Jan 2025
AI
In a First, AI Models Analyze Language As Well As a Human Expert
www.quantamagazine.org/in-a-first-ai-models-analyze-language-as-well-as-a-human-expert-20251031/
274 sats
\
0 comments
\
@0xbitcoiner
31 Oct 2025
AI
Vibe physics
www.math.columbia.edu/~woit/wordpress/?p=15012
2355 sats
\
4 comments
\
@south_korea_ln
1 Aug 2025
science
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open LLMs
arxiv.org/abs/2402.03300
762 sats
\
0 comments
\
@zuspotirko
6 Feb 2024
science
Controversial math proof divides experts in bitter academic dispute
www.earth.com/news/mochizuki-controversial-math-proof-divides-experts-in-bitter-academic-dispute/
201 sats
\
2 comments
\
@south_korea_ln
14 Jun 2025
science
To Have Machines Make Math Proofs, Turn Them Into a Puzzle
www.quantamagazine.org/to-have-machines-make-math-proofs-turn-them-into-a-puzzle-20251110/
268 sats
\
0 comments
\
@0xbitcoiner
11 Nov 2025
AI
Computer Scientists Combine Two ‘Beautiful’ Proof Methods
www.quantamagazine.org/computer-scientists-combine-two-beautiful-proof-methods-20241004/
398 sats
\
0 comments
\
@0xbitcoiner
6 Oct 2024
science
more