Forgot your password?
typodupeerror
AI China United States

Should US Open-Weight AI Labs 'Distill' Frontier Models Too? (techcrunch.com) 108

Silicon Valley giants and national security experts "are calling for action against Chinese companies engaged in model distillation," reports CNBC. But "I would do nothing," says Y Combinator CEO Garry Tan. "We could argue that there should be an American distillation regime." Distillation is the process of using the outputs of a more capable AI model to train a smaller or less capable one, sometimes illicitly... Anthropic has accused Chinese companies such as Moonshot AI, DeepSeek, and MiniMax of the practice, while OpenAI believes DeepSeek's V3 and R1 model architectures were distilled from its own GPT-4 and GPT-4o models. In the midst of this, the U.S.'s National Security Agency, Cybersecurity and Infrastructure Security Agency, and Federal Bureau of Investigation released an official cyber security advisory warning on the topic on Tuesday...

But Tan believes regulators should focus less on curbing distillation and more on creating an equilibrium between open weight models and frontier models — as long as frontier models retain a price premium that allows their business model to remain feasible. "This is actually the ideal case. You want open weight models to give people freedom and access," he explained. "If I were a regulator, that's what I would go after." Tan acknowledged that this is a hard balance to strike, calling it "a tightrope." Nevertheless, he says it's a balance worth pursuing — saying it "could result in the best possible outcome."

Tan later told TechCrunch he'd like to see America with more open-weight options that aren't Chinese, built by smaller U.S. open-weight AI labs using those same training techniques on products from America's frontier AI labs: Anthropic CEO Dario Amodei had previously publicly called on U.S. regulators to crack down on distillation. It's notable that the commander of Silicon Valley's prestigious and prolific startup accelerator doesn't agree.

To be clear, Tan isn't advocating for American AI labs to use stolen credentials to distill. He wants them to be free to come in the front door. In fact, his argument is twofold. He feels it's an overreach for AI labs to dictate what their customers can do with the information their models share with them. He also notes that the proprietary AI labs didn't ask permission when they vacuumed up as much human knowledge as they could to train their models. They famously ingested plenty of copyrighted material without the permission of those intellectual property holders. "Controlling what users and customers do with API calls to closed weight models feels constraining, and there's a role government can play here to normalize the fact that access to intelligence that was trained on broad public access data should itself also be more a form of a public good than something locked away behind restrictive terms of service," he told TechCrunch when asked why American labs should be free to distill, too...

To him, the true AI doomer scenario is for all the immense power of frontier AI to wind up in the hands of a single powerful, proprietary provider. "The nightmare scenario, the doomer scenario for AI is that there's just one company," he said. "It has the best access to capital. It has the best AI researchers. It runs away with it and suddenly there's one company that's monolithic. And that would be bad."

This discussion has been archived. No new comments can be posted.

Should US Open-Weight AI Labs 'Distill' Frontier Models Too?

Comments Filter:
  • by Njovich ( 553857 ) on Monday September 14, 2026 @03:09AM (#66336940)

    I suggest that banning distillation can only be considered if Dario, Sam, Sundar and Elon make sure they negotiate deals and pay for all copyrighted material that they used for training models in the past.

    Lets just take the minimum statutory damages for copyright infringement, $700 per work.

    Still interested Dario? Dario? Where did you go?

    • by fluffernutter ( 1411889 ) on Monday September 14, 2026 @03:21AM (#66336944)
      I think we should just have the same laws for everyone. That's never going to happen though.
      • Re: (Score:2, Interesting)

        by Anonymous Coward

        To be clear, the same enforcement, punishments and justice.

        Right now the AI companies can hack other companies and nobody goes to prison, they can practically treat it as marketing cost.

        In contrast if you or I did the same thing, we are likely to end up in prison.

        But yeah like you said, never gonna happen. In most of the world there's a Caste system. And the Untouchables are at the top.

        That's why Sony, Amazon, Microsoft, AI bros can all do unauthorized modification of computer systems and not do any prison

      • by bmajik ( 96670 ) <matt@mattevans.org> on Monday September 14, 2026 @10:32AM (#66337220) Journal

        I used to agree with you, and then oddly enough, about 25-30 years ago, I came across this signature on another regular slashdot posters postings:

        "The law, in its majestic equality, forbids the rich as well as the poor to sleep under bridges, to beg in the streets, and to steal bread."

        (AI tells me the attribution is "The Red Lily", Anatole France, 1894)

      • by dfghjk ( 711126 )

        Is that what you were thinking when you supported Trump's election?

    • by jhoegl ( 638955 ) on Monday September 14, 2026 @03:28AM (#66336950)
      I believe restricting data is the answer to make a much better model. "AI", is just a search and compile engine. It doesnt think like people assume, it doesnt rationalize, it doesnt even understand what it is looking at.
      For example, researching RNA, if you "teach" it the basics, and see if it "learns" on its own to get itself to where we are today, it wont know what to do. It cant come up with a theory and then logic through how to validate.
      It CAN however, link likenesses together after learning how to. And the less slop it has to deal with, the more efficient and accurate it can be.

      Anthropic is just trying to buff up its IPO release, which is likely this year or early next year.
      • by gtall ( 79522 ) on Monday September 14, 2026 @04:06AM (#66336960)

        More to the point, Anthopic is trying to kneecap the Chinese models so that American frontier model companies can bask in warm glow of their exorbitant pricing.

        • by jhoegl ( 638955 )
          Now that makes more sense.
        • Local models are becoming more and more capable on less and less hardware. This should actually really scare the Frontier modelers and it sounds like it is. I can't wait to see what I can run next year on my hardware.

          Dell actually sells this really impressive $10,000 AI server computer thing with 128gb of vram. Consider what I can accomplish with 8gb vram and considering how stupidly priced a single nvidia 32gb graphics card cost, that $10,000 a full server with 128gb vram sounds like a fair ask. Assuming y

      • by Rei ( 128717 ) on Monday September 14, 2026 @11:43AM (#66337304) Homepage

        It doesnt think

        Yeah, it does. [transformer-circuits.pub]

        it doesnt rationalize

        Yeah, it does. [transformer-circuits.pub]

        These are not Markov chains. They're neural nets. They work via extremely complex chained fuzzy logic on superpositions of conceptual states.

        And the less slop it has to deal with

        This is literally a thread about distillation, aka, training on the outputs of other models. Synthetic data is the cornerstone of modern training. "Model collapse" is not something that actually happens in the real world, only in contrived settings, the model equivalent of if you could lock a person alone in a dark room with only their thoughts for ten thousand years.

        • Wow thanks for sharing those links. I'll have to dive deeper into them when I'm not at work, but I have to say those nodes / supernodes popups are fascinating!
        • > It doesnt think

          >> Yeah, it does. https://transformer-circuits.p... [transformer-circuits.pub]

          ?

          In the paper, in the text itself, whenever the authors use the word think to refer to a human thing they don't put think in quotes and when they use the word to refer a description of what an LLM is doing they always put "think" in quotes.

          I didn't find an explanation why. Are they in effect acknowledging that they don't mean the same thing by the word in each situation and want to avoid a kind of equivocation? Near the end of the

          • by Rei ( 128717 )

            Yeah, I used to do that too. Decided to stop bothering with the quotation marks a couple months ago.

            We're not going to spend the rest of our lives putting quotations around words when talking about models. "Think" and "reason" the words we have in English for what is going on. No need to tiptoe around it. Again: models are not humans. They are not the same as us. But those are the words we have in English for what they're doing.

            • models are not humans. They are not the same as us. But those are the words we have in English for what they're doing.

              There are no doubt better words, and English has never been shy about using loan words, so there's no excuse for contaminating those words by using them for something they're not suited to.

              • by Rei ( 128717 )

                Why invent a new "aithink" verb when we already have "think"?

                • Why invent a new "aithink" verb when we already have "think"?

                  Because what it's doing is clearly not the same thing, and also because we have other words which are better suited to the case like "process". If you're not going to invent a new word, or borrow a word, then at least use the word which makes the most sense.

                  • by Rei ( 128717 )

                    Because what it's doing is clearly not the same thing,

                    Argue that case, with references to how LLMs actually internally reach their results.

                    The physical biology is certainly different, but this isn't a question about "what things are made of" or even the specific NN type (e.g. smooth vs. spiking), and training differences don't even come into the picture; it's a question of the broad strokes of how conclusions are reached on forward processing.

                    • > it's a question of the broad strokes of how conclusions are reached on forward processing.

                      I think it's uncontroversial that in the wide sense of the word think, what the brain does, cognition, or neural processes, is not at all like what computers do including LLMs.

                      Forward processing, in my opinion, doesn't change that. The computer processes involved can be referred to as thinking but it is only in an anthropomorphic sense that doesn't actually transfer over to human thinking except maybe as a conve

        • These are not Markov chains. They're neural nets. They work via extremely complex chained fuzzy logic on superpositions of conceptual states.

          They are, by definition, Markov chains. It's a Von Neumann architectural Finite State Machine processing a string of bytes, outputting bytes that get fed back in. That's a Markov chain.

          Don't get hung up on bullshit about "extremely complex blah blah blah.". That's just designed to confuse you.

          • by Rei ( 128717 )

            They are, by definition, Markov chains.

            Even in your attempt to be pedantic here (in which the universe and everything within it is a Markov chain), no, it's not. The hardware state is Markovian but the linguistic processing is a Nth order autoregressive process; it depends on the N previous states. Also, your argument is akin to saying "a Boeing 747 is just an arrangement of quarks, so don't get hung up on aerodynamics." it entirely ignores the relevant architectural details, and instead substitutes a m

            • We've had this conversation in excruciating detail before, about 6-9 months ago, and we've already established that you simply don't have the knowledge about Markov chains to have a meaningful discussion on LLMs. You're just wrong about this, and I'm not going to repeat myself.
              • by Rei ( 128717 )

                Hey AI, who is being more reasonable in this conversation?

                User 2:50PM
                Who is being more reasonable in this conversation?

                [Snip]

                Model 2:50PM
                ThinkingThoughts
                Expand to view model thoughts

                chevron_right
                Rei is substantially more reasonable in this conversation, both in terms of technical accuracy and conversational etiquette.
                ere is a breakdown of why:

                1. Technical Accuracy and Explanatory Value

                martin-boundary’s argument relies on vacuous reductionism:
                martin-boundary claims that because an LLM runs on a digital

                • This is why you believe this bullshit. You're abdicating your thinking to the machine and then sucking down what it gives you.

                  • by Rei ( 128717 )

                    And meanwhile, you keep presenting nothing more than absurd pedantic deflection from the means that LLMs actually use to achieve tasks, trying to mislead people into thinking that they're just disguised probability tables, ignoring the actual consequences of your derailment of the conversation from actual mechanisms to an exponentially-exploding model of the consequences of said actual mechanism, and even in your pedantism, failing to understand the difference between a Markovian state (the physical hardwar

    • by Rei ( 128717 )

      I kind of wonder if the best anti-distillation strategy is, if you detect suspicious traffic from someone (which happens a lot, they monitor for anything that looks like distillation), instead of blocking them, feed them say the output from Llama 3.1 8B or whatnot ;) Maybe finetune it a bit so it talks Claude-ish. But basically, subtly poison their dataset with hallucinations and crappy reasoning without it being immediately visibly obvious.

      As for copyvio, sorry, this is something for the courts, and so f

      • I imagine Anthropic etc are doing behavioral analysis/segmentation on all user accounts to detect bots and other interesting things, because why wouldn't they, and the distilling botnets are tuned to fit acceptable organic-appearing behavioral segmentation. Don't look random, look meaningful. Don't be a super-power user, but be productive enough not to set off any alerts.
        I was always wary (maybe in hindsight) of the free data thing. It goes both ways as we have seen.

    • by McLoud ( 92118 )

      I suggest that banning distillation can only be considered if Dario, Sam, Sundar and Elon make sure they negotiate deals and pay for all copyrighted material that they used for training models in the past.

      Lets just take the minimum statutory damages for copyright infringement, $700 per work.

      Still interested Dario? Dario? Where did you go?

      Too little too late, the Chinese et all will continue doing their thing and it seems their models are already pretty capable, its't just a matter of time until "frontier models" gets surpassed.

  • by locater16 ( 2326718 ) on Monday September 14, 2026 @03:28AM (#66336948)
    "Frontier" models trained on everyone else's data, according to the companies that made them it's perfectly legal. Thus it's perfectly legal to train off those same models, obeying terms of service and robots.txt is as perfectly optional for US open models as it is for these supposedly trillion dollar companies. Train away until it's actually as open source and free as in Linux and I never have to hear about Sam fucking Altman ever again.
    • by OrangAsm ( 678078 ) on Monday September 14, 2026 @04:49AM (#66336982)
      Distillation on the frontier... sure sounds like making moonshine out in the woods... illegal, but it'll be done anyway. Might as well accept that some inbreds will be distillin'
    • by quenda ( 644621 ) on Monday September 14, 2026 @06:22AM (#66337008)

      "Frontier" models trained on everyone else's data, according to the companies that made them it's perfectly legal.

      No, few things are perfect. But the basic principle is that you cannot own ideas or information.
      Intellectual property laws protect secrets, trademarks, copyright and patents.
      Copyright relates to the expression of an idea, not the idea itself. The output of a pre-trained AI might violate copyright, but the training itself does not.

      Thus it's perfectly legal to train off those same models,

      There is no "thus". Are you expecting the law to be less than 10 years behind reality?
      What has happened is that the distillers have made use of the frontier models in ways that violate the "Terms of Service". You know, those documents that you never even read?

      • That gives Anthropic the right to suspend service.

        Thus its perfectly legal to train on your exfiltrated data.

      • No, few things are perfect. But the basic principle is that you cannot own ideas or information

        You don't know what you're talking about. Owning ideas is precisely the point of patents.

      • by DarkOx ( 621550 )

        I think there is room to argue though, that distillation might violate copyright while training on large volumes of large material does not.

        At that point you are modeling a model. I think you could argue that it is a bit like copying a novel where you hand type it; adding lots of typographical errors a long the way; get laze and leave out a chapter or two that you don't feel moves the story a long much, then slap a new title on a call it your own.

        I don't think any court anywhere would say did not violate t

        • by ceoyoyo ( 59147 )

          I think there is room to argue though, that distillation might violate copyright while training on large volumes of large material does not.

          Nope. There is an argument about whether training a model on source material violates copyright or not. Current court decisions in the US say that the training does not, but you must have acquired the material legally in the first place.

          Current court decisions in the US say that the raw output of a model is not copyrightable. Never mind that the people doing the distill

          • by DarkOx ( 621550 )

            Ultimately it won't matter either way. Open source models will get better. They don't even need to achieve absolute parity with the frontier models they just need to get to 'good enough for most tasks most normies need to do'.

            Some of them are probably already knocking on the door of that, and where they fail the frontier models often do as well.

            Even if GenAI has dramatically increasing your business productivity, $100 a month per white collar workers becomes a pretty big expense item on the balance sheet p

            • by ceoyoyo ( 59147 )

              It matters quite a lot if someone sues you. Or if a propagandized public and their leaders get into a war over it.

      • by drinkypoo ( 153816 ) <drink@hyperlogos.org> on Monday September 14, 2026 @10:30AM (#66337212) Homepage Journal

        What has happened is that the distillers have made use of the frontier models in ways that violate the "Terms of Service".

        The question at hand is whether the terms of service should even be allowed to contain those clauses. If you're paying for access to the model, it should be your business what you use the output for. If they don't want you using the data you got out of it, they can not sell you access to it in the first place.

        • It does seem a bit weird.

          When you pay for API usage, are you buying the model's output or are you licensing it under some restricted usage terms?

          It seems what companies like Anthropic are saying is "you are buying the output, but we won't sell it to you unless you agree not to use it to compete with us".

          It'd be like Microsoft saying we'll only sell you a C compiler if you agree not to use it to write a compiler.

      • The Chinese also use grey market intermediaries that operate transit hubs using stolen/hacked credentials to access and distill the frontier models, so their massive-scale distillation attacks go quite a bit beyond a ToS violation.
        • It's hard to get upset when China is stealing from a handful of companies that themselves stole stuff to get where they are today. Just makes them all look like deceitful assholes. There are no good guys here.

          • I am certainly unhappy that everything I've ever written online at StackExchange, Reddit, the open web, and Slashdot has been used to train AI models, which in turn have been distilled by China to train their own competing AI models; so that they can automate the jobs we've done and leave our children without the same routes to employment. In China's case, it's also to try and take over global hegemony from the United States. Ordinarily I'd think eating into US hegemony would be a good thing for the world,
        • The Chinese also use grey market intermediaries that operate transit hubs using stolen/hacked credentials to access and distill the frontier models

          This is an oxymoron. A market that peddles in stolen shit isn't a grey market it is a black market.

      • Thus if they don't care about the law, I don't care about the law. You don't get to break the law to build your business then turn around and tell everyone, "but you have to follow the law" Yeah, fuck you, I won't.

        If they don't play by the rules, why should we? I say, distill away everyone. At the rate local models are coming along, we won't need these big companies anyway. You'll just run if off open source free as in beer software.

  • by BitterEpic ( 10503015 ) on Monday September 14, 2026 @04:11AM (#66336966) Homepage

    The business model of AI companies is to steal copyrighted material as well as private data. The argument against distilling makes no sense to me. You have taught your AI drones this is the new Napster era payed for by Wall Street.

    • The argument[s] against distilling makes no sense to me.

      They're all motivated reasoning, you shouldn't take them too seriously.

      That is, they start with the conclusion they want, then say whatever they think will confuse/convince you the most.

  • ... I know how my parents felt when I tried to explain programming too them. It all makes grammatical sense but the actual meaning eludes me.

    Is it just me? Am I just getting old and heading for the get-off-my-lawn stage or are there any other people who like to think they're tech savvy but are totally confused by AI models and how they really work?

    • The vast majority of people have neither the theoretical background, nor the necessary attention span for understanding how an AI model works. You can catch up on the theory if you have the attention span. My parents, for example, never got around to even using computers. My uncle, however, did. He did not become a programer, but he could use a computer somewhat comfortably, although he did fall victim to fraudsters once.
      So, keep reading, look up wikipedia and youtube for learning, and keep your pins priva
      • by Viol8 ( 599362 )

        Thanks for the heads up, I never thought of searching for the info. Been there, done that, I still don't understand whats going on on a large scale inside these models.

      • by ceoyoyo ( 59147 )

        Bullshit. The fundamentals of modern AI models can be understood by anyone with a basic knowledge of algebra, which you should have picked up in junior high. They're piecewise linear approximations and use exactly the same equation as the linear regression you learned in high school or first year university. The more advanced stuff is hacky restrictions on that basic design to tone down the model's flexibility and make it easier to fit.

        The reason it's hard to understand is because a) people who have no idea

    • I was just confused by the gibberish of this headline.

    • by Jeremi ( 14640 )

      It's really the same problem that we see with quantum physics -- the subject matter is esoteric enough that common sense can't be used to determine whether what you're listening to is really advanced theory or utter BS (or some combination of the two); they both sound the same, and the situation isn't helped by the presence of large numbers of people who think they know what they are talking about but don't. :/

    • How they work, and/or how they are trained?

      Anything specifically?

      Distillation is the word used to refer to the practice of using the outputs of one LLM as inputs to train another, often smaller one.

      The big US companies like OpenAI and Anthropic have usage "terms of service" that forbid you from doing this, but Gary Tan is saying this is unreasonable and you should be allowed to use the outputs of an LLM in any way you choose.

      Specifically, Tan is hoping that, if allowed to, some US companies will choose to d

      • Considering the output of these models can't be copyright, I don't see how they have any legal standing to say you can't use it how you see fit. People are even paying to get that output, so boo to anthropic for people using the output they paid for. That's more then anthropic did to get started.

        • I think what they are, in effect at least, trying to say is that you do own (not license) the model output, but as a pre-condition for using their service you agree not to use the output to compete against them.

          It's basically as if Microsoft said you can't buy our compiler if you are going to use it to build a compiler (or a clippy, or anything else that we do).

    • If you have a 8gb nvidia card, you can educate yourself very easily on AI models. Here are a few links that talk about and explain some of this. It should be enough to get you started.

      https://techtactician.com/begi... [techtactician.com]
      https://roblaughter.medium.com... [medium.com]
      https://stable-diffusion-art.c... [stable-diffusion-art.com]
      https://github.com/msb-msb/awe... [github.com]

      This last link is to a Dell computer that I thought would be pretty sweet if I was really into local AI a lot more. I mentioned it in couple other posts but here's the actual link.

      https://www.dell [dell.com]

  • by kasnol ( 210803 ) on Monday September 14, 2026 @06:40AM (#66337018) Homepage

    This feels like the PlayStation ownership drama all over again. What exactly do we own anymore?

    If I pay for an AI service and its answers and chat logs are stored on my own SSD, can I use those outputs to train my own model? If not, what exactly did I buy?

    When I create a Word document, Microsoft does not suddenly own it or call it a “distilled Microsoft document” because Word helped produce it. So why should an AI provider automatically retain control over what customers create with its service?

    Maybe the deeper problem is that some companies are still thinking in an old ownership model, where they assume they can own every layer forever: the platform, the intelligence, the outputs, and even what customers subsequently learn from those outputs.

    That gets especially awkward when US AI companies are simultaneously defending lawsuits over copyrighted training data while arguing that competitors learning from their model outputs is unacceptable. Meanwhile, geopolitically, some companies that are not bound by these restrictions have an ugly incentive to exploit that advantage.

    Perhaps distillation is not the real disruption. The disruption is discovering that once intelligence becomes a service, controlling every downstream use of it becomes a giant grey area: who owns what [X], and what can [X] legally be used to do [Y]?

    • Yeah it feels weird. It's like microsoft going after someone that uses Visual Studio to make their own IDE. As long as the Chinese customers are paying for access to the models than this seems like entirely fair game. I'm also rather confused as supposedly Kimi K3 was distilled from Fable/Mythos but they only had about 10 days of access to do that? If it's possible to learn the secret recipe to an AI model in that short of a time, it hardly seems like a secret.
      • It's like Apple being mad a Microsoft for stealing "their" idea for the desktop even though Apple got it from Xerox. There's a great scene at the end of Pirates of Silicone Valley that sums that up nicely and more elegantly then I did.

  • Pot meet kettle (Score:5, Insightful)

    by bloodhawk ( 813939 ) on Monday September 14, 2026 @06:44AM (#66337020)
    Kind of amusing, All the AI models steal IP from artists the web, books all without paying and claim fair use. Now when it is done to them it is unfair.
    • by bmajik ( 96670 )

      I think you're repeating something that feels emotionally valid but in my experience, isn't factually (or legally) accurate.

      I have found AI models to be frustratingly _obedient_ of copyright.

      I'll be in a situation where I ask a specific question about what a document says, and the AI repeatedly says the work is copyrighted and it can't just tell me what it says _verbatim_, but it can answer specific questions for me in its own words (while telling me to check the original).

      Two examples from my life are proc

      • There is a huge difference between regurgitating copyrighted material to you and training with it. The former is clearly copyright infringement, while the later is what they actually do as it falls in the grey space.
        • by bmajik ( 96670 )

          I don't think training a model with copyrighted data is any kind of gray zone at all.

          Running inference against such a model _might_ be infringing.

          But I don't see the legal argument that model training would be.

          Think of model training as lossy data compression. You provide an input, it produces a compressed artifact, and it can no longer reproduce the original when you uncompress it. It can produce _something_. It cannot reproduce the original.

          On slashdot, we all agree that format shifting copyrighted wor

          • Don't get me wrong I think it is completely fine to train the AI that way, I just think it is purely hypocritical to then claim it is unfair when others use the output in a similiar way.
          • You forgot about derivative works.

            The definition is pretty fuzzy and varies for different types of media but the general standard is recognizable elements were reproduced.

            No court has yet quite taken that up that I know of but you don't need to reproduce a work in full to have a derivative work which could be a copyright violation.

  • by xack ( 5304745 ) on Monday September 14, 2026 @07:31AM (#66337038)
    You go to school to "distil" the teacher's expertise, you distil from the books and websites you read. The actual main problem with distilling from an another AI is model collapse due to recursive slop training. With all human knowledge sources already scraped, they only have got themselves to train on.
    • by gweihir ( 88907 )

      Maybe simplistic humans. Well, to be fair, that covers most of them.

    • You see this collapse in people in cultures (religious, isolated, etc) where they only study received knowledge, only "what the ancestors knew", etc, and do not research or innovate themselves. After a while, they start acting like a poorly anchored LLM and less like an actual thinking creature.

      • by ceoyoyo ( 59147 )

        You see it in the way you just used the word "you." Model "collapse" is a misnomer. It's drift, which you see in humans in literally everything they do, and you can demonstrate to yourself just by repeatedly generating random numbers and calculating the mean.

        • Randomness is the problem in LLMs that leads to inconsistency of hallucination. It's being substituted for the part of the processing we do that software can't do yet. Hardware can't do it either, only wetware. The part that (inductively? Instinctively?) knows whether the output makes sense just plain doesn't exist in LLMs.

          • by ceoyoyo ( 59147 )

            Ah, you're one of the people who thinks that your neurons are gateways to the unphysical realm? Neurons in general, or the pineal in particular?

            Also, this has nothing to do with my post. Did you reply to the wrong one?

    • Yes, but the models are getting smaller and better. If you have never played with local models, you just don't understand why Anthropic and co are worried. The local models are making big strides. Each year they get better and more efficient. I can't wait to see what I can run next year compared to what I can run today.

  • by gweihir ( 88907 ) on Monday September 14, 2026 @08:57AM (#66337106)

    As LLMs are based on a massive theft of information, everybody should now steal from the LLMs as well. It is just fair.

  • by nicolaiplum ( 169077 ) on Monday September 14, 2026 @09:03AM (#66337110)

    We're back to Vizzini's claim in The Princess Bride: You are trying to [take] what I have rightfully stolen

  • We need a book that explains how AI technology functions.

    There is a huge amount of news and comments, but apparently nothing that helps people understand how AI is developed, and functions.
    • by Sloppy ( 14984 )

      I thought "Mommy Has Four Graphics Cards" was pretty good, but "Not Everyone Computes" did a great job of explaining why the industry is largely about remote services rather than running it locally.

    • by ceoyoyo ( 59147 )

      A little tutorial if you know a little bit of Python:

      https://numpy.org/numpy-tutori... [numpy.org]

      And a book with more detail if you're more into math:

      https://www.deeplearningbook.o... [deeplearningbook.org]

      Everything else is pretty much scaling up and introducing some restrictions on the basic model.

  • It's hilarious to me that the owners of these "frontier" LLMs complain about others "illicitly" training on their models, when their models were trained mostly illicitly on the work of others.

  • Did they pay for the usage?

    Did your AI give them the responses they paid for?

    Not really seeing what the problem is here, other than throwing a piss fit over someone using your product successfully.

  • The irony is almost too perfect. AI companies have spent years vacuuming up the entire public internet to train frontier models. Then, when someone systematically queries their models’ outputs to train smaller models (i.e., distillation), suddenly it’s “illicit,” “theft,” “industrial-scale attacks,” and a national-security issue requiring FBI/NSA advisories.
  • This solves a lot of problems:

    - Since AI can't be copyrighted, its later use shouldn't be prohibited. You can use the results of the AI you paid for, for whatever, including distillation. Trademark law still applies, ToS against abuse still applies, there's just no preventing people from taking it and using it for something competing.
    - Distillation makes open weights more powerful, competing with the frontier labs.
    - The frontier labs thus won't make as much money, which means building thousands of datacente

  • "Should US Open-Weight AI Labs 'Distill' Frontier Models Too"

    What the fuck does that headline even mean? Reading it makes me feel like I had a stroke.

    • The headline is questioning what Gary Tan is advocating for, namely that:

      1) The US frontier labs shouldn't be allowed to impose terms of service that restrict how users can use the model's outputs

      AND

      2) Some US labs should take advantage of 1) to distill a, presumably smaller, model from the outputs of these frontier ones, and should release the resulting model in open weights form

    • From what I know off the bat: there are two types of "AI" (being "Large Language Models"): the "frontier" models, which are trained over human-made stuff: work, arts, behavior, etc., and the "distilled" models, which uses other AI's to train them, which is often a much cheaper thing to do.
      It's something like: frontier models imitate humans, distilled models imitate these imitators.
      What's generated by a training of a LLM are these datasets called "weights". Like code, a company can keep it closed or just ope

      • by CAIMLAS ( 41445 )

        Nope, they're all "distilled" models. That's what post training is, largely, and they're all using other models to help curate the datasets for this. They're all doing it. It's why all the models appear to be converging in terms of capabilities and personality: they're effectively using the same feed stock and using an increasingly similar set of techniques.

        US "frontier" labs are distilling, as well. It's postulated (with good reason) that Opus 5 is, for instance, just a distillation of Fable to a smaller m

  • The scores of Chinese models like Kimi K3 are pretty much the same as the scores of the US frontier models - even if you could somehow prevent "distilling from US models", that would mean nothing for the rest of the world. And since several Chinese models are open-weight already, there are no horses left any AI company could keep in their barn. (Even if those horses were not stolen from others, before.)
  • They say "sometimes illicitly"... OK, so what's the criteria that makes it illicit?

    If they can't state that, it's just FUD.

    And frankly, that's most of what the anti-AI brigade is spewing: it's a psyop, pure and simple - by both China and the US domestic companies, with the blessing of the US government. People who can't see this aren't paying any attention to what's happening.

    Anthropic/OpenAI want restrictions to lock people in and provide them with an out for their gross overspending.
    US government and its

  • "Distilling" a model does not result in a better model.

    The only point in it is: you can automate it.
    But still you have the problem of validation.

    Sure, you can distill dozens of models and try to feed all that into one model.

    A law against it, does not make any sense. You can not enforce it, you can not prove the fact that it got broken.

    What is the point?

  • The arrogance is insane. What a fucking waste of time to argue this when your business model is exactly the fucking same.

A memorandum is written not to inform the reader, but to protect the writer. -- Dean Acheson

Working...