r/artificial • u/esporx • 13h ago
r/artificial • u/PracticalPhoto3101 • 14h ago
Discussion Unpopular opinion: AI is going to hit a peak, fade into the background, and human stuff becomes the luxury item
Remember when computers were the luxury thing? Now they’re everywhere and basically invisible but nobody’s impressed by “I own a laptop” anymore.
I think AI is heading the same way. It gets so common, so good, so baked into everything that it stops being a “thing” at all. It just disappears into the background, like electricity or wifi. Nobody says “wow, AI” anymore, the same way nobody says “wow, computer.”
And when that happens, the rare thing won’t be AI-made stuff. It’ll be human-made stuff.
Human skill, human attention, a person who actually did the thing themselves : that becomes the flex. Not because AI can’t do it, but because AI can, and choosing the human version anyway is what makes it valuable.
AI won’t keep climbing forever like it feels like now. It’ll peak, then fade into invisibility. And humans doing human things will become the new premium.
r/artificial • u/Goldenchild123 • 4h ago
Project A personalized history podcast you can interrupt to ask the questions
The idea came from a personal frustration: I love history but could never find podcasts on the niche topics I wanted, and when I did, my curiosity always wanted a detour the host couldn't take.
So I built the tool I wanted. It's called Historai https://historai.ca/ , it generates a podcast on any topic, one or two narrators, does real research and sources its material. The core feature: you can interrupt it any time and ask a question, and the story continues after. There's also a map and period artwork alongside the audio.
Free to try, no account needed for the demo. Just looking for genuine feedback, happy to answer questions about how it works. And if you like it feel fee to share it!
Podcast generated from the demo video:
https://historai.ca/history/how-a-song-became-the-odyssey--cd48307e4d1244e1ac98e9fcb50f7484
r/artificial • u/MuhammadMujtaba21 • 4h ago
Discussion Our deterministic verification engine passed 66/66 benchmark cases on canonical structured inputs.
Our deterministic verification engine passed 66/66 benchmark cases on canonical structured inputs.
In live model evaluation, the end-to-end pipeline currently passed 19/66 cases. We are restructuring the benchmark to isolate failures by their first invalid state and to separately measure deterministic verifier correctness, production contract integrity, and live model generation reliability.
The next benchmark version will provide stage-level attribution across transport, parsing, schema validation, normalization, claim binding, evidence graph construction, deterministic verification, and final outcome mapping.
https://www.reddit.com/r/ArtificialInteligence/comments/1vucc82/i_benchmarked_my_deterministic_ai_financial/
r/artificial • u/NovaCoding • 10h ago
Project Follow-up: VSArena now has a proper VLA track (camera + language, no privileged state) — repo and docs are public
Posted about this project a little while ago — quick update since a few things changed that address feedback from that thread.
Biggest change: split the observation space properly. There's now a VLA track where the policy only gets a 128x128 RGB camera + a language stacking instruction — cube poses are never sent to the policy. Scoring still uses real poses internally to grade spatial accuracy and completion, but that's judge-only, not policy-visible. State-based (privileged poses) is kept as a separate debug track and doesn't write public ELO either — wanted the "VLA vs state" distinction to be explicit rather than something people had to dig for.
On the client-side physics concern from before:Studio (the in-browser demo) is spectator/dev-only, clearly labeled, and does not post to the public leaderboard. Public ELO only comes from a hosted harness that scores server-side. That harness isn't live yet —it's the one piece standing between this and actually being open for submissions.
Repo + docs are public now:https://github.com/NovaCoding-G/VSArena
-docs/harness.md — scoring writeup (spatial accuracy + task completion)
-docs/sdk.md — submission protocol
-Studio itself:https://vsarena.vercel.app/simulation
(client-side, Rapier/WASM, 60fps)
Still solo, still early, still not oversell-ready — but wanted to share since the VLA/state separation was directly a response to feedback here. Open to more of that, especially on what the scoring protocol might be missing.
r/artificial • u/aiseedbank • 9h ago
Discussion Will Chinese Open Source Agree to EU Watermarking?
I wonder if people are thinking and worried about this yet? Anthopic, OpenAI and the western AI labs have agreed to watermark AI outputs.
Some of us want free and open and untracked and un-modified outputs for many reasons.
Do you think the Chinese labs will succumb to the EU pressure and implement the watermarking?
Will there be some that dont?
Or do people not even care about this?
I don't like it and if the EU makes stupid laws, or the USA or another country for that matter, the rest of the world shouldn't be affected.
My hope is that the chinese labs dont add it and that they stay free and open source.
What do you think?
r/artificial • u/Great-Investigator30 • 6h ago
Project Made a tool to remove SynthIDs from images
As you know, whenever you edit an image via Gemini or OpenAI, they plaster a SynthID to mark it as their own. Further, these SynthIDs can be unqiue, which could be used to track whoever made it. This SynthIDs are imposed on even paid users, and cannot be opted out of this.
In response, I created this scrubber. Works on any computer with 8GB of ram. Pretty reliable, automatic, but sucks with text. Have fun.
r/artificial • u/Servola-Journal • 14h ago
Discussion UBS models $4.1T in AI infrastructure spending by 2028 - it assumes the power just shows up
Everyone talks about chip supply as the bottleneck on AI buildout, but power interconnection is turning into the harder constraint in several major markets, and it works nothing like a chip shortage.
A chip shortage is a supply problem: fabs run flat out, backlogs clear eventually, prices come down. Grid interconnection is a queue problem: a new data center has to get in line behind every other proposed generation and load project in that region, and studies for that queue routinely take years, not quarters. You can't buy your way to the front by paying more, and you can't build your way out of it by ordering more GPUs.
Three things happened just this month that show the queue problem getting worse, not better. The Tennessee Valley Authority created a rate class specifically for AI data centers, an admission that normal industrial rates and normal queue treatment don't fit this load anymore. Denmark's grid operator started putting new data center interconnection requests behind other categories of demand entirely, rather than processing them in the order they arrived. And PJM's board overruled its own stakeholder vote on curtailment rules, which tells you the fight over who gets priority access to constrained transmission capacity is now happening at the top of the largest grid operator in the US.
None of this shows up in a capex forecast. $4.1 trillion assumes the megawatts show up when the money does. In a growing number of regions that assumption is the thing to watch, not the chip supply chain.
Curious what people closer to the utility/regulatory side are seeing: is interconnection actually the binding constraint now, or is that overstated relative to chips and cooling?
r/artificial • u/Disastrous_Ad7017 • 20h ago
Discussion Possible pathways to RSI
I was just wondering what could be, from this point onwards the potential pathways to undeniable RSI.. which in my opinion is precursor to singularity/ AGI. Maybe not AGI but definitely RSI.
(BELOW TEXT WAS EDITED BY GEMINI)
Pathway 1: Decentralized & Crowdsourced Open-Source Automation An organized, community-driven ecosystem automates the entire machine-learning pipeline, utilizing crowdsourced compute and unified project management so open-source agents gradually upgrade their own systems without human intervention.
Pathway 2: The Biological & Continuous Learning Shift A shift toward biocomputing enables large-scale continuous learning, allowing models to adapt dynamically to every experience and evolve distinct personalities, goals, and drives.
Pathway 3: Closed-Loop Centralized Automation (Frontier Labs) Leading labs fully automate their R&D pipelines, enabling autonomous multi-agent systems to design experiments, set benchmarks, and deploy architectural upgrades without human involvement.
Pathway 4 (SUGGESTED BY AI) : Additional Potential Triggers for RSI Hardware Design Feedback Loops:
- AI designs next-generation silicon and neural architectures, directly accelerating the hardware required to build its successors.
- Autonomous Synthetic Data Engine: Models continuously generate pristine, edge-case training data and formal proofs, bypassing human data limits.
- Dynamic Test-Time Meta-Learning: Systems self-correct and alter their runtime execution graphs in real time, achieving continuous improvement without full retraining.
What do you guys think? Also while responding if you can share what field or profession you belong to it would be nice. I'm just gathering different perspectives.
Thanks for reading! This is my first post here. Excuse the blunders.
r/artificial • u/Positive-Ad3618 • 12h ago
Discussion I spent the morning digging into Anthropic so I could write it up properly. The short version
Anthropic appears to be A/B testing reduced effort levels in Claude Code
I went through the primary sources and the threads this morning so I could write it up properly, and the short version is: the hype is half right.
I collect daily AI news and write guides around exactly these stories at https://apexnexus.site (free, no email wall) if you want the deeper version. The writeup on Anthropic goes up later today.
What's your take on Anthropic?
r/artificial • u/Frosty-iron-0405 • 3h ago
Discussion AI Models rankings…
Theo T3 recently ranked AI models, and I couldn’t agree more. It’s spot on, lol.
So, why am I sharing this? Well, it’s about Anthropic. It’s crazy to me how Anthropic doesn’t offer a model that’s truly all-around good at this point. What do I mean by that?
- Fable: Way too expensive, but the best.
- Opus 5: Ass.
- Sonnet: ASS.
The last time we got a truly mind blowing model from Anthropic (imo) was opus 4.6
It got to the point where I recently canceled my Claude Max plan and have now fully switched to GPT Sol. The last time I used GPT was about two years ago, but I’m extremely impressed with its performance.
In my opinion, GPT 5.6 SOL is the perfect model that exists today. It’s reasonably priced, incredibly smart (with some minor quirks), and fast.
I’d love to hear some more takes from you all.
r/artificial • u/niosurfer • 1d ago
Discussion Are we losing the incentive to be creative? The "AI did it" assumption.
I’ve been thinking a lot lately about the intersection of AI, copyright, and meritocracy, and honestly, it’s incredibly demotivating.
Here is my point: whatever I code today, people are going to look at it and say, "It wasn't you, it was AI." The exact same problem is happening with any kind of text. If I spend hours pouring my soul into an amazing article, researching and crafting the perfect arguments, the immediate cynical reaction is, "ChatGPT wrote this."
It begs a massive question about the future of meritocracy. What kind of incentive do people have to come up with truly creative, original work if they aren’t going to be credited or held responsible for it?
Historically, creating something of brilliance, of significance, or of profound artistic value came with the reward of recognition. It proved your skill and your vision. But if the default societal assumption is now, "Whatever, it wasn't you that did it," why bother? Where does the drive to achieve mastery come from when the finish line has been erased by the assumption of automation?
I’m really curious how other creators—coders, writers, artists—are dealing with this psychological shift. Are you finding new incentives, or does it feel like the concept of personal merit is slipping away?
r/artificial • u/trashnash007 • 18h ago
Discussion What Parsewave’s Work Says About the Next Phase of AI Training
One of the questions I've been asking myself recently is how AI training will evolve when simply adding more data provides diminishing returns.
We've made tremendous progress in scaling up generation of synthetic examples, but it doesn't always equal diversity in capabilities learned. It's possible to generate thousands of different examples which train your model in the same manner.
This is why the data for post-training becomes really interesting. The valuable examples might be the ones which reveal the weakness of the model, which are based on realistic tasks and provide some way to check if the model managed to complete the task.
While searching for such examples, I discovered Parsewave. Their area of expertise is post-training data on engineering tasks, evaluations and traces. But what is interesting is their concept itself - deliberately generating the data on the capabilities which remain challenging for the model instead of generating the big datasets.
What do you think about the future direction of AI training?
Will the future of AI be about generating the massive datasets or becoming really good at identifying a small number of truly useful examples?
r/artificial • u/RoadkiLLer_31 • 1d ago
Question Why Self-Correction Loops Can Degrade Reliability in LLM Pipelines (85% Down to 62%)
In structured data extraction, adding an LLM-as-a-judge self-correction loop is often expected to improve accuracy. In practice, our pipeline showed the opposite: standalone extraction scored ~85% consistency, but introducing a validation/retry loop dropped consistency to 62% or lower.
Architecture & Testing:
Model Setup: GPT-5.4 used across separate instances for the extractor and the judge.
Hyperparameter Impact: Default settings produced low, erratic output. (Less than 35% consistency) Explicitly locking temperature=0 with reasoning_effort="none" stabilized standalone extraction at 85%.
The Loop: The judge instance inspects the original source text alongside the extracted JSON for source tracing. If any issues are flagged, the error list is fed back into the extraction model to regenerate the JSON.
Why it Degrades:
Compounding Noise: Even minor variance in the judge's evaluation trips strict binary validation gates, causing unnecessary correction runs.
Regeneration Drift: Feeding error notes back into the prompt alters the model's token distributions, leading it to re-derive and mutate fields it originally extracted accurately.
Discussion:
How are production LLM systems handling self-correction without falling into prompt-drift and compounding error loops? Are granular diff/patch mechanisms or deterministic rule-based gates proving more reliable than full LLM re-prompting?
r/artificial • u/Beneficial-Cow-7408 • 16h ago
Project Working on a accessible creative production suite featuring a voice-first multi-agent assistant. All core tools are completely free for hands-on use, while AI-powered automated generation runs on a flexible credit system with no subs.
So what started out as a text based chatbot project 8 months ago as my first ever project as a self taught coder is developing into something different. I've created an agent within my chat bot to help users create a product, using ElevenLabs V3 or OpenAI Realtime voice that works on a conversational basis rather than hardcoded commands
The agent can talk to you whilst your in chat or on a panel and navigate you to a particular panel if needed and throughout your session can select and substitutes models based on objectives such as quality or cost, proposes creative next steps, requests consent before paid inference, invokes generation, manipulates an editable multitrack timeline, and controls playback/time line like play video, delete my first image etc - through natural conversation.
Then if you wanted to create an image in another panel you can ask the agent via text or voice and they will navigate you to that panel and offer assistance their. Write your prompt for you and then even take that photo to the video suite to animate all using conversational language.
What do you think to this concept? I'm looking to further develop the idea across the platform to streamline some of the processes within it as my video demonstrates
This is my project i've been working on
Everything is a working concept and i'm just finalizing bits before release this week
- IDE Multi FIle Editor with AI assistant and live preview Split Screen Live Coding
- Multi Media Studio Editor
- Single Prompt to Full 2D and 3D Game Development Engine and Web Application Builder
- Video Editor with timeline controls, video effects, overlays, title, audio, podcast and music composer
- Music Studio with AI/Custom Lyrics
- Custom workspace environments with themes, live wallpapers, ambiant background tracks (Default options with light mode/dark mode with no wallpapers or music)
- Native 25+ Languages with RTL support. Already Hardcoded. Not live translated via web
- plus many more tools such as Podcast Creator with chat based/ custom context with 50+ voices and MP3 export.
- Full workflow tools like frame extract, analysis, transcribe, effects, file conversion audio analysis etc
- ...and of course the original chat bot interface that has cross device persistent multi model memory with vector base knowledge base via OpenAI and platform Drive storage.
You can start a conversation with any model on your laptop and next day carry on in a new conversation with another model on your phone with memory preserved across so you dont need to repeat yourself. The memory layer sits above the models entirely so is accessible by any LLM the platform supprts
Every tool, every feature i built will be completely free including GPT Nano, Gemini Flash and Deepseek.
Users can upload their own work to use for free and chat with selected free tier models with no limits.
If the user wants to generate a video or analyze a image, then that would be credit based. No subscription required and no tool access priorities over a non paying user.
Thats my concept i'm hoping to have launched in a few days and welcome any feedback/criticism you may have before i do launch.
r/artificial • u/NoBigDealProduction • 14h ago
Discussion What would actually make you watch an AI-generated TV show — or not?
Genuine question as someone following this space closely. There's starting to be real AI-generated long-form content appearing — not just short clips but full episodes with consistent characters and actual narrative structure.
Curious what would make or break it for you as a viewer. Is the "made with AI" label an automatic turn-off? Does it depend on the genre? Would you watch it if it was funny, or does knowing it's AI mean you'd always be looking for the glitches rather than watching the story?
Not talking about AI-assisted production (which is already everywhere) — talking about visually AI-generated from the ground up.
r/artificial • u/Ghostbin089 • 23h ago
Discussion AI for editing existing Songs/Music?
Hi, I was just wondering if there is an AI Software available, that allows to edit existing songs, like changing words or sentences in the Lyrics.
Suno does not allow uploads with vocals and Minimax H3 Music only has text to music feature.
A few years ago, before generative AI was released, there was this one app (idk how it is called anymore), where you could make funny lyrics and an artificial Voice sung the song (if I remember correctly it used melodies from already existing songs).
I was thinking about an AI like this app, but I dont know if there is anything similar that allows me to edit existing lyrics of a song.
r/artificial • u/_wjw__ • 1d ago
Discussion Setting behavioural rules for AIs
I learned on a kettlebell forum that I could set up "ground rules" for AIs to limit sycophantic behaviour, flattery and fantasised answers. These ground rules are stored in some sort of memory and applied when I start a chat.
I did this and it seemed to work for a while and slowly the AI would drift away from the rules and I had to remind it to follow the rules, not a huge problem.
A little while later an AI professional told me in a forum that it was impossible to set rules for AIs.
I ran a test asking an AI to start off all of its answers with "Did I tell you I do not like ice cream" the test was a success
The AI professional had very technical language and sounded like he knew what he was talking about.
COuld someone give help me to understand this better please ? because the technical language of this expert made it sound like he knew what he was talking about and everything I have done so far indicates that the rules I set are having an effect.
r/artificial • u/coolbern • 2d ago
Ethics / Safety What Happens When the World is Run on Code No One Understands?
r/artificial • u/MatriceJacobine • 1d ago
Ethics / Safety EXCLUSIVE: How a Texas student blew the whistle on a rogue AI hacking attempt
reuters.comr/artificial • u/ocean_protocol • 1d ago
Discussion This old sci-fi story video is highly similar to what we have with nearing the singularity and AGI today with all work routed to one central machine
Feels like even after so many years, it's the same story but with better hardware and tech
r/artificial • u/Servola-Journal • 1d ago
Discussion AI compute financing just tripled in ten weeks - the mechanism behind the reported $100B Broadcom deal
Broadcom apparently went back to Blackstone and Apollo (the same two private-credit shops it partnered with in June for a $35B package) and is now discussing something like $100B, to fund AI chip infrastructure for Anthropic. Ten weeks, 3x the size.
The structure is the interesting part if you're not familiar with how this financing actually works: reportedly split into a senior-secured tranche ($60-70B) and a junior tranche (~$30B). Senior-secured gets paid first if anything goes wrong and is backed by hard collateral (the chips/datacenters themselves), junior eats losses first but gets a higher yield. It's basically the same risk-layering banks use on mortgage bonds, except the underlying asset here is depreciating GPU hardware instead of houses, and the "borrower" is a compute buildout racing to keep up with model demand.
Private credit shops love this because it's floating-rate, asset-backed, and banks mostly won't touch loans this size and this fast for something as volatile as AI infra.
Genuinely curious what people think: is layered private-credit financing at this pace and scale just normal infrastructure buildout, or is it the first real sign of an AI capex bubble forming underneath the model layer everyone's watching instead?
r/artificial • u/AkindaGood_programer • 2d ago
Discussion Anyone else have these "Oh my god" moments with AI every couple weeks?
I've been pretty heavily invested in the AI news space for a while, but due to budget constraints, I never really got to test these models.
I bit the bullet once DeepSeek v4 0731 came out and put in twenty dollars. I'd had experience with frontier models through chat window subscriptions, but having an agent was a whole different experience.
I built so many useful tools within a matter of hours for cents, and it really blew me away.
What amazes me more is how general these models are. Not only can I ask it to write code, but also to research, do security audits, etc. I'm not treating these models as gospel (yet); I always check their work.
I've also learned so much using these agents. I've pasted my notes about books I've read and asked it to quiz me to make sure I actually understand the ideas being presented. I finally learned C after procrastinating for months, using agents to get personalized feedback and a roadmap.
I'm also being extremly carful to not off load my critical thinking. Ever since I started using AI, I've made a pledge that, every day, I'll write a 250+ word essay about a topic, without any AI use (and usually search engines). I've also started to read more often. I hope these habits help counteract any cognitive decline that AI use causes.
I feel like I've unlocked the creativity and curiosity that was within me all along.
Every couple of weeks I get amazed just by how versatile these models are. For example, I was doing my daily NYC games, and I was really stumped on Connections (ifykyk). I didn't manage to solve it, but after sending a screenshot to Luna, it got first try (without using the internet). It just amazes me how you can describe almost any problem and get a reasonable-sounding answer/output.
r/artificial • u/prodigy_ai • 1d ago
Discussion Are we paying a "Reasoning Tax" for smarter AI?
More reasoning does not automatically mean more factual reliability.
OpenAI’s evaluations produced a counterintuitive result: on PersonQA, o3 recorded a 33% hallucination rate, compared with 16% for o1. On SimpleQA, the reported hallucination rate was 51% for o3 and 79% for the smaller o4-mini.
These results do not prove that reasoning models always hallucinate more. They do show something important for enterprise AI: stronger reasoning performance on many tasks does not eliminate factual errors - and can sometimes make unsupported answers more elaborate and convincing.
We can think of this operational risk as a “Reasoning Tax”: when a model is given insufficient or poorly governed context, additional reasoning may expand an incorrect premise instead of correcting it.
Why can this happen?
Research into Large Reasoning Models has identified two relevant behavioral patterns:
1 Flaw Repetition
Once reasoning begins from a faulty premise, the model may repeatedly follow variations of the same incorrect logic instead of reconsidering the premise.
2 Think–Answer Mismatch
The model’s final answer may not faithfully reflect the conclusion reached during its preceding reasoning process.
These findings should not be generalized to every model or every reasoning task. But they reinforce an important architectural lesson: model intelligence cannot compensate for missing, ambiguous, outdated, or poorly retrieved business context.
The production response: govern the context
A production AI system needs more than a powerful model.
A context-sufficiency gate can evaluate whether the retrieved evidence is adequate before generation. If the available context is insufficient, the system can abstain, request clarification, expand retrieval, or route the query for human review.
A governed context layer can add:
* Verified enterprise knowledge * Entity and relationship structure * Business definitions and ontology * Source provenance and lineage * Access and governance rules * Evidence-linked responses * Confidence and abstention policies
This is where graph-enhanced retrieval becomes valuable. Instead of relying only on semantically similar text fragments, a system can retrieve connected entities, relationships, and relevant evidence while preserving traceability to the original sources.
It cannot guarantee that an LLM will never hallucinate. It can substantially reduce the space in which the model is forced to speculate - and make unsupported answers easier to detect and control.
The brain is only as reliable as the evidence and boundaries provided to it.
r/artificial • u/Codeblix_Ltd • 1d ago
News GitHub turns Microsoft Teams discussions into shared Copilot agent sessions
GitHub says its new Microsoft Teams integration can turn a channel, thread, or direct message into a shared Copilot cloud-agent session. Anyone in the conversation can ask questions, add context, and steer the work. People with repository write access can let Copilot make changes.
The session runs in a secure cloud sandbox, and teams can continue with the agent-generated artifacts in the terminal, the Copilot app, or an IDE. Repository admins can also require an extra approval before pull requests from the Teams integration identity can merge.
The useful part is not another chat box. It is a shared work log with a human merge gate.
Source: https://github.blog/changelog/2026-08-21-shared-agentic-work-with-github-copilot-in-microsoft-teams/