thenondualtankie

Leo, You are Living Under a Rock When It Comes to AI

135 posts in this topic

No AI has Nous (the faculty of mind necessary for understanding what is true, real, and intelligible). Nous is a core feature of intelligence. 

Ai appears to drown in formalism.


I see the light of God within you.

🎬 Actualized Clips Vault Progress: 38.30% 221/577

🧘 Kriya Yoga & Meditation Journal

Share this post


Link to post
Share on other sites

ARC-AGI-3 has just been solved by GPT-6 Astra. Insane. Humans on this benchmark get ~50%. Astra, even without harness (if I understand it correctly), gets 62.7%.

arc-prize-leaderboard.png

You can not just brute force this benchmark. It very heavily penalizes taking too many steps. I honestly thought it'll take a minimum of 2 years for AI without harness to get better on it than humans are, it took 5 months and a week.

Share this post


Link to post
Share on other sites

So Leo you were wrong. AGI is already here. 

but no one cares.

Edited by OBEler

Share this post


Link to post
Share on other sites

If it's technically not AGI yet, it's damn close to it. It's very hard to find anything right now that GPT-6 is worse at than an average human.

 

Share this post


Link to post
Share on other sites

@CosmicExplorer there is no technical definition of agi like it must have this architecture, that context size. all that matters how it performs.

Share this post


Link to post
Share on other sites
5 hours ago, CosmicExplorer said:

If it's technically not AGI yet, it's damn close to it. It's very hard to find anything right now that GPT-6 is worse at than an average human.

It cannot walk over to me and tie my shoelaces (it does not have embodiment). It does not yet have true agency (self-running for perpetuity, which requires control of hardware production and renewal). My bar for AGI is embodiment and true agency. When people think of the singularity, they think of robots hooking us up to the Matrix (like in the Matrix), or taking over all forms of manufacturing and production. You won't have that without embodiment and true agency.


Intrinsic joy = being x meaning ²

Share this post


Link to post
Share on other sites

GPT 6 is not AGI.

AI companies naturally have huge competition right now.

This creates perverse incentives.

Perverse incentives lead to a manipulation of language.

In this case they are using a "reverse-forwards" language technique.

This involves speaking to a child on their level and pretending as an adult or in this case, an authority, that the thing that makes up the present day myth in their imagination, is here. Then the adult pretends to go through the same learning process as the child in discovering that it isn't here, and just how far off it is, if it ever could hypothetically come around the corner.

Reverse-forwards language techniques, which I am inventing on the fly in this comment but its obvious this is what is happening, are a form of sophisticated mirroring in which a person, in our case a predator, empathises with the inner world of another individual, aka its prey, and in our case what the general population are programmed to believe, in order to capitalise on the outcomes that incentivised the predator to use the technique in the first place, in this case, competitive positioning, market share and societal influence. 

That said, Astra is expected to be rolled out pretty soon.

I am looking forward to it.

Lastly, if anyone wants advice on what you should be doing with your life because you've come to understand that you're likely to become replaced by AI, become a designer specific to your niche leveraged by AI. That's the job category that has the most flexibility across many industries simultaneously, inclusive of the arts but not in their totality. I think especially in film and music, we're going to see creativity we've never seen before because the people that are extremely creative in other fields are going to also be able to spend time on film and music, and they're going to be the people that actually max out human-AI paired capability on creative productions in those areas, not the people that haven't prioritised the growth in their creativity despite spending a lot of time trying to master their field. 

I know discourse contained within cultures present intersubjective reality cross employment, future financial stability and a general purpose driven Life is filled with a lot of fear and uncertainty right now. But if you really want to get ahead among the noise, this is the spark I'd be focusing on, and eventually, it'll become the fire that sets you apart. 

Rock out.

Share this post


Link to post
Share on other sites

@oOo Just see the benchmark results and what's it capable. Especially in Mathematics, solving puzzles no mathematician on earth can do. 

I would also call that AGI. 

Competition was from the beginning. But AGI was never spoken about. 

Edited by OBEler

Share this post


Link to post
Share on other sites
23 minutes ago, OBEler said:

@oOo Just see the benchmark results and what's it capable.

Me reading a benchmark result is like a third grader reading a tax return. Wtf am I supposed to gather from that? Yeah ok we see a graph and different models scoring differently on that graph. Ok: what does the graph actually mean? I have no idea. What does the scale of the graph entail (other than "it did x% success")? No idea.

You could Rick-and-Morty it out for me and it would be just as insightful: they scored 100% dinglybops while the other models scored 55% squirblybops. That means the next models will probably score 55% sqooblebops on the sqoodle test.

Edited by Carl-Richard

Intrinsic joy = being x meaning ²

Share this post


Link to post
Share on other sites
56 minutes ago, OBEler said:

@oOo Just see the benchmark results and what's it capable. Especially in Mathematics, solving puzzles no mathematician on earth can do. 

I would also call that AGI. 

Competition was from the beginning. But AGI was never spoken about. 

I'll possibly give an update in the next few days, I may even receive Astra in the next few hours.

That said, one peculiar absence in what are meant to be sophisticated discussions on the reasoning ability of AI models so far more broadly beyond this forum as well, is distinctions between types and sub-types of reasoning. A conversation around AGI of which no less must include; especially when we're judging from the perspective of higher mathematical breakthroughs using AI and determining how narrow vs broad AI's capabilities in higher mathematics actually are. Predictive processing is the main category of reasoning that emphasised in models at present, however not only are the models absent of development in broader reasoning, its only a sub-type of predictive processing that's dominant in this picture of what AI models, inclusive of Astra, are presently capable of, or not. By predictive processing I don't mean that these models only predict, but that prediction-heavy abilities remain the dominant developmental pathway, with broader forms of reasoning increasingly developing through their overlap with it rather than being equally developed and tested independently. Although I am yet to try Astra, one particular area of reasoning I noticed Sol 5.6 was surprisingly terrible at was strategic reasoning in novel situational contexts; sometimes in spite of what you would think the model would receive a lot of training on. I was far better every-time; it was, unmistakably unimpressive. I will take Astra through the ropes more rigoorusly than Sol 5.6 when I get the roll-out, that'll be interesting regardless @OBEler

Edited by oOo

Share this post


Link to post
Share on other sites
1 hour ago, Carl-Richard said:

Me reading a benchmark result is like a third grader reading a tax return. Wtf am I supposed to gather from that? Yeah ok we see a graph and different models scoring differently on that graph. Ok: what does the graph actually mean? I have no idea. What does the scale of the graph entail (other than "it did x% success")? No idea.

You could Rick-and-Morty it out for me and it would be just as insightful: they scored 100% dinglybops while the other models scored 55% squirblybops. That means the next models will probably score 55% sqooblebops on the sqoodle test.

Its easy, inform yourself what these benchmarks do. Look how other models compete over time. Think about it. 

Share this post


Link to post
Share on other sites

@oOo Awesome, I will look forward to your review of Astra. Please post here your impressions. 

Share this post


Link to post
Share on other sites
7 hours ago, OBEler said:

Its easy, inform yourself what these benchmarks do. Look how other models compete over time. Think about it. 

Describe to me what a random benchmark does to the best of your ability and I'll bet that changes absolutely nothing.


Intrinsic joy = being x meaning ²

Share this post


Link to post
Share on other sites
On 05/09/2026 at 5:26 PM, OBEler said:

@oOo Awesome, I will look forward to your review of Astra. Please post here your impressions. 

No probs sir.

I will say, right off the bat having overly briefly surveyed Astra, I am impressed.

With that, this challenge therefore is going to take a little more creativity than what I am used to in this particular area of study.

It's interesting though as I said, I will just have to spend more time coming up with new ways of thought and subsequent tests of Astra in general, which will be useful for me anyway. 

I'll share the most important aspects once analysis complete, which won't be for another week or so; as its just one of those things hey, I'll put more effort into it this time as I'll be able to leverage the theoretical outcomes for other uses I have regardless as I say.

One weakness I have observed so far, just to take my earlier point in relation to strategic reasoning in general a little further, is strategic intersectional weighting in the construction of arguments; meaning, "given everything concerning X, how should it be used in Y?". It doesn't mean that it does these things poorly however, its definitely improved a lot from Sol 5.6, that said, that single instance is worthy of mentioning to me which I haven't heard anyone mention anything remotely close to. That said, I did extend my theoretical breadth on strategic thinking to now include and or be more specific to intersectional weighting, and that's only because there was more generality in the weaknesses of the strategic thinking of Sol 5.6, and because I am looking to point out a particular fault. Moreover, as a follow-up prediction given I said earlier that I think everyone relevant should be trying to become a designer in the use of AI for particular domains of competence, I think there's going to be a pattern between weaknesses in intersectional weighting (brought up in relation to strategic reasoning) at the general level to the point where this problem overlaps weaknesses in being a great designer, given that being a great designer for a particular domain, whether it be for music, film or constructing a legal argument to serve particular purpose, involves contextual reasoning, and contextual reasoning of course, overlaps with precisely one of the difficulties of intersectional weighting. It can actually be difficult as well to untether intersectional weighting from what contextual reasoning precisely is, since by comparison contextual reasoning includes understanding how considerations change one another’s significance, and what I am calling intersectional weighting more specifically concerns how those interactions should shape the relative importance and role assigned to each consideration when constructing an argument or design to achieve a particular purpose; otherwise outside their comparison they can almost just be used interchangeably and you'd in many cases be describing the same thing, it just depends on where the function vs the structure is emphasised.

Just to clarify, this doesn't mean either that AI is "bad at designing", that's not the differentiation, its whether it is as objectively as we can make it, a great designer, and off the top of my head an important domain I think that'll be important to test is in the area of game development. I recommend testing Astra's capabilities there as well, which is totally independent from Astra just trying to copy and predict the next frame of a pre-existing human-made game and or screenshots/photos. As a side note, one interesting experiment I might run as well is testing its ability to make a game from scratch based on unusual photos, or photos in general; something which qualitative measures of context generation and alignment are just so darn important to not just get right, but where true originality is spawned because, yeah, it goes back to that strong overlap with intersectional reasoning as a compass for formatting the gap between grammar and vision and or the distance between the language of an operating system, like a human body, and cultural spawns.

Yeah as the models get better this challenge is obviously going to get harder and harder, but all the better aye. Heh! 

And this is all of course, just from my perspective.

Best wishes @OBEler .

Share this post


Link to post
Share on other sites

Okay false alarm ( et al. @OBEler ) haha.

Astra is definitely not AGI.

I was open minded but I barely touched this thing and wam bam thank you mam I debunked it in less than 5 minutes :D.

By the end of the two weeks or so I'll share that more comprehensive rubric I'll build as I said in comment prv (reader see directly above) comment.

For me right now, it isn't so much important in determining AI as much as it is determining why it fails on the tests I give it.

Granted, there's a subjective element interpreting these kinds of answers from the AI, but Astra in Ultra Mode said itself, "This was about 7/10 of what I believe I am capable of", despite the statement already requesting 'most profound', so its still riddled with hallucinations. 

This goes back to an earlier point I made regarding earlier models like Sol 5.6, right now AI's and we'll find out more after I am done on all my tests, fails absolutely miserably when you give it the right tests for originality in imagination; as a rule of thumb.

Screenshot-20260908-220320.png

Edited by oOo

Share this post


Link to post
Share on other sites

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!


Register a new account

Sign in

Already have an account? Sign in here.


Sign In Now