Rendered at 12:46:54 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
tristanj 2 days ago [-]
Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data.
I enabled all data sharing settings but still don’t have a message about free use on that screen - the help page says free tokens are available to “some” users - is that 1% of users, 40% of users, etc?
Does your screen have the message that you’re getting free tokens?
I had no idea about this before. I just enabled it.
You're eligible for free daily usage on traffic shared with OpenAI.
Up to 250 thousand tokens per day across gpt-5.4, gpt-5.2, gpt-5.1, gpt-5.1-codex, gpt-5, gpt-5-codex, gpt-5-chat-latest, gpt-4.1, gpt-4o, o1, and o3
Up to 2.5 million tokens per day across gpt-5.4-mini, gpt-5.4-nano, gpt-5.1-codex-mini, gpt-5-mini, gpt-5-nano, gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, o1-mini, o3-mini, o4-mini, and codex-mini-latest.
Usage beyond these limits, as well as usage for other models, will be billed at standard rates. Some limitations apply. Learn more.
moonu 20 hours ago [-]
Yeah, I have it enabled for just one of my projects, and it says "You're enrolled for complimentary daily tokens." Haven't been billed for any usage with this. Their tooltip doesn't mention the newer models, but it works for those too. I used them in a recent project that I knew wouldn't hit the rate limits. Not sure how eligibility is decided though.
OsrsNeedsf2P 1 days ago [-]
I tried following this page, and it's certainly a lot more complex than what Meta is offering. Different price tiers, opt-in configurations, usage based availability.. I'll take the 10x discount for flipping a param switch over this all day long.
stingraycharles 1 days ago [-]
Seems pretty clear to me: enable it for the projects you want, and there’s a 1M / 10M token limit per day, depending on the model you use. Assuming an average context size of 100k tokens, that is 10 to 100 requests, which is not a lot. Reason enough to prefer actually paying for Meta as well.
no-name-here 1 days ago [-]
Does it show the free tokens message for you after you enable sharing? It does not for me.
Also, in a desktop browser at the page's [1] lower left it says "Personal organization", so if the 'organization' term is the concern, OpenAI still seems to use the term 'organization' even for personal accounts.
Subscription plans share data by default (it is possible to opt out though)
ignoramous 1 days ago [-]
In my experience using DeepSeek v4 Flash free tier (context size limited to ~200k), the model isn't nearly as good as the paid one for agentic coding tasks. Unsure, if that's the case with these other providers too, though I wouldn't be surprised if it indeed is.
markasoftware 22 hours ago [-]
There is no official dsv4 flash free tier api afaik. Did you get it from open router or something? Likely quantized.
The paid official dsv4 api already shares data with deepseek for training
jjcm 2 days ago [-]
I actually really like that pricing strategy. It's very transparent
rvz 23 hours ago [-]
Developers on this orange site never ever learn.
p-e-w 23 hours ago [-]
What are you talking about? This is one of the most honest offers ever made by a corporation. “We’ll use your data, and we’ll compensate you for it.” Where is the problem?
drob518 21 hours ago [-]
I agree. I don’t mind you training on my data if you’re up front about it and you are willing to compensate me for it in some way. And if I don’t like the deal, I can always pay full price or use a different model.
RideOnTime22 21 hours ago [-]
Because they want to play teams and act like there's "good" companies,
all the while they throw their bank account and whatever they want at GPT or Claude models.
bdangubic 2 days ago [-]
I love the idea of this pricing strategy but there is no way meta is not training on your data regardless of your monthly invoice
simonw 1 days ago [-]
So you think the only difference between the $1.25/million token plan and the $0.10/million token plan is that you pay them more to both lie to you and breach their contractual obligation to you?
customguy 18 hours ago [-]
"they 'trust me'. dumb fucks."
The only difference between students doing it then and professionals doing it now is the students had no positive, glaring reason to mistrust.
SirLordBoss 17 hours ago [-]
I can't understand the psychology of people who think like this.
You really think a throwaway quote when Zuckerberg was a college student applies nowadays?
You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what? Middling amounts of training data?
lunar_mycroft 6 hours ago [-]
Meta literally ran a man in the middle attack to spy on it's user's encrypted network traffic when they used third party apps [0]. More recently (and relevant to this issue), the engaged in industrial scale piracy to get training data for their LLMs [1]. The idea that they have changed since Zuck was a college student creeping on his female classmates and now wouldn't commit actual crimes against their own users in order to get a bit more data is just demonstrably false.
> You really think a throwaway quote when Zuckerberg was a college student applies nowadays?
What does "nowadays" mean? What changed?
> You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what?
And what is the risk? Food companies are at risk when they put addictive chemicals into food because they get controlled regularly. What is the equivalent here?
bdangubic 1 days ago [-]
when has Meta ever not broken their contractual obligations (I am being serious here)? are we seriously discussing/expecting any sort of privacy related to Meta?
you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...
JimDabell 1 days ago [-]
> when has Meta ever not broken their contractual obligations (I am being serious here)?
If you are being serious, then you have a wildly distorted view of the world. No organisation can routinely break all of their contractual obligations. If you think Meta are doing this then you are not seeing Meta, you are seeing a fictional bogeyman.
simonw 1 days ago [-]
This is one of those situations where we are both completely baffled by the position held by the other person.
The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.
ben_w 1 days ago [-]
> The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.
To me, their history suggests that they know more than most about how to get away with breaking both the spirit and the letter of the rules, and that they are motivated to enrich themselves without regard for what rules are broken.
I would not know which to expect, spirit or letter, in any given instance.
However, even if they were to surprise me by being perfectly meticulous about the letter of the rules from now on, I have so little trust in them that I would expect some technicality somewhere in the language of the contract.
seanp2k2 20 hours ago [-]
This. Without a massive cultural shake-up and turnover of upper management, why would expect them to behave any differently when they've been rewarded so heavily for this behavior in the past? Meta is ultimately an advertising and data brokerage company -- they make money selling and leveraging user data and behavior and are "bound" by fiduciary duty.
Honestly, I wish the tech community would do a better job identifying the actual individuals who are making these decisions instead of associating them with the brand they're under at the moment, because it's not THAT many people. Like, if you look at only the 100 tech sector companies included in the $NDXT index, how many individuals hold a VP or above title there, and how difficult would it be to trace key decisions at different times made at different companies to the individuals holding those positions there plus board membership and major shareholder identities (with the caveat of known unknowns here) and make a sort of ethical index and trace that along their careers with company moves, promotions, board appointments, shareholder decisions, etc? Go a step further and link that to financial performance and I'm sure folks at quant firms are already ten steps ahead of where I'm going with this, but I care less about profiting off of this data and more about surfacing it to show that it's people and more specifically, specific individuals driving these decisions.
teiferer 1 days ago [-]
How does money change that trust? They certainly have breached their word on this in the past (for non-paying users of Facebook).
simonw 1 days ago [-]
Because they didn't have a financially backed contract with those non-paying users.
customguy 18 hours ago [-]
does the contract enable the customer to monitor/search Meta to ensure they are honoring the contract? if there is no mechanism for that it means very little. Though I bet/hope some will feed them "watermarked"/unique but worthless things and watch for traces of that to pop up in models or something like that, but that's hatdly enough to just take their word for it.
simonw 15 hours ago [-]
That tends to be what discovery in lawsuits is for.
customguy 12 hours ago [-]
That's circular reasoning, how would there be a lawsuit if customer have no way of knowing?
simonw 10 hours ago [-]
Sensible companies don't take that risk.
lunar_mycroft 6 hours ago [-]
Sensible companies don't do a lot of stuff Meta has already been caught doing - sometimes with real consequences (but never enough to actually deter them, of course) - in the pursuit of more data.
ignoramous 20 hours ago [-]
This is what the Muse Code launch blog post says:
We're also beginning to accept requests for zero data retention. Contact Meta sales to request this.
Zero data retention and "we don't train models on your input" are different things.
Retaining data is common for investigating abuse.
bdangubic 1 days ago [-]
I respsct the F out out you and all your work but I am completely baffled by this line of thinking, we are talking about Meta here…
simonw 1 days ago [-]
I just cannot comprehend this level of corporate villainy that boils down to:
Let's have two pricing levels. One will be 12.5x more expensive than the other. For the cheaper one, we will get their express permission to train models on their inputs. For the more expensive one, we will still treat their data EXACTLY the same, but we'll lie to them and say that we won't. I've checked in with legal and they raised their sherry glasses and toasted "Gentlemen, TO CRIME".
trollbridge 23 hours ago [-]
Because data a user doesn’t want you to train is probably much better data?
sgt101 22 hours ago [-]
What I love is the idea that this is more likely to happen at Meta than at Anthropic, Google, Microsoft or OpenAI... or in the PRC!
There's contractual cover, there are lawyers all over the USA ready to make themselves very rich by creating a class action over it... you don't have to worry! Or, you do, but frankly, only MI6/CIA/Mossad can save you now.
ignoramous 22 hours ago [-]
> I just cannot comprehend this level of corporate villainy that boils down to:
Easy to comprehend: Trust lost is hard to gain.
Though, it is extremely competitive of Meta to sell Muse Spark (Grok 4.5 / Qwen 3.8 Max level model) cheaper than DeepSeek v4 Flash / MiMo v2.5 / GPT 5.6 Luna, regardless.
bdangubic 21 hours ago [-]
Again, this Meta we are talking about.......
We can make this fun, within 18 months from now, there will be some story / whistleblower like "a inadvertent defect was found that allowed your 'private' data to be used in our training endeavours, we sincerely apologize and have already addressed the issue" - if this does not happened in this timeframe I will donate $1k to a charity of your choice.
malshe 23 hours ago [-]
Knowing Meta anything is possible. It’s not like they never shafted their paid customers. They have been overcharging advertisers by showing wrong metrics for years. I’m sure at some point they will come out with raised hands and admit to this “glitch” they found during an internal review.
arolihas 22 hours ago [-]
No one has to say it and plan it out loud. But the data will be sitting there. The incentive to improve the model for enterprise use will only get stronger. It doesn't take much for one engineer or team to go rogue to hack a benchmark. There was a whole cheating controversy with llama 4.
j_maffe 21 hours ago [-]
What does cheating benchmarks have to do with breaching financial user agreements?
To use your argument: all it takes is one whistleblower to get the company sued for billions of dollars.
arolihas 10 hours ago [-]
What does unethical behavior have to do with unethical behavior? What does looking at data you're not supposed to look at have to do with looking at data you're not supposed look at? Please don't be obtuse.
People whistleblow on Meta all the time. Most recently they got fined half a billion for suppressing child safety research. Kids getting groomed. It doesn't make a difference to a company that has a net profit of $60 billion a year.
ray_kay777 2 days ago [-]
This makes it a very interesting alternative to Deepseek for personal work where I don't care about the training - judging by the AA benchmarks it seems like overall cost per task is similar to the new Deepseek Flash but with better benchmarks (and inbuilt vision capabilities).
sgt101 22 hours ago [-]
Do you know a good place for accessing the full sized chinese models but hosted not in China?
drob518 21 hours ago [-]
For Deepseek specifically, OpenRouter has a number of non-Chinese providers. Pricing varies.
martinald 1 days ago [-]
Also looks incredibly fast. 150tps on openrouter (nearly all deepseek providers are around the 50tps mark).
embedding-shape 2 days ago [-]
Except with one you know they'll release the weights and architecture back to the community, with the other, it leans towards they won't do that.
drob518 20 hours ago [-]
Does anyone know if this is available via OpenRouter or just directly via Meta? I've looked on OpenRouter and it just shows the standard pricing with a single provider. Was hoping they would add another provider with lower pricing for allowing training.
dan15 2 days ago [-]
Probably to compete with DeepSeek, which AFAIK also retains data (or at least OpenRouter says they do)
dghlsakjg 2 days ago [-]
FYI: there are providers of deepseek that offer the same or lower pricing and zero retention policies.
greyb 2 days ago [-]
Unfortunately, none with the same caching performance as DeepSeek proper.
solenoid0937 1 days ago [-]
I don't trust them, they are probably distilling on your data.
dghlsakjg 14 hours ago [-]
It would be pretty wild if cloudflare or digital ocean opened themselves up to a massive lawsuit by doing this.
Do you have any evidence that cloud providers that don’t train their own models are violating their own contracts and risking their own good reputations by stealing data?
ForHackernews 2 days ago [-]
DeepSeek is really crazy cheap, though, and they don't have a giant pool of other invasive personal data to correlate it with.
jofzar 2 days ago [-]
I hate to say this and this is because I fucking despise meta. But between DeepSeek and Meta, and trust they handle the training data correctly, I trust meta.
aand16 2 days ago [-]
What do you mean by "correctly"?
handfuloflight 1 days ago [-]
For example, OpenCode says they have a ZDR with DeepSeek. Some of us are skeptical that's going to be properly honored. There's no way to know.
tpm 1 days ago [-]
Right about DeepSeek, but with their history of handling data, I would never trust Meta either.
HDBaseT 2 days ago [-]
Meta, please offer this on OpenRouter too (ZDR + Non-ZDR, official Meta Provider).
GodelNumbering 2 days ago [-]
I think that's a fair offering tbh
theanonymousone 1 days ago [-]
I believe that will put them in the Pareto frontier.
But I cannot find this "variant" in OpenRouter.
Havoc 17 hours ago [-]
Always a bit surprised by this. 10x is a sizable discount. And as fun as my crappy projects are I struggle to see it being of much value as training data
6thbit 19 hours ago [-]
So what Meta believes fair for "paying" for your data is $0.1/Mtok plus the opportunity cost of $3/Mtok in output?
Jumpstylish 24 hours ago [-]
SOLD! Mr. Zuckerberg, I’m a dumb f—- and I trust you!
deno 2 days ago [-]
I think it's limited to US or at least EU is excluded.
Instantnoodl 1 days ago [-]
Just noticed it too... Seems like I wasted my time setting up a account to test with the discounted pricing
ray_kay777 1 days ago [-]
Yep, doesn't work in Australia either.
discordance 1 days ago [-]
Call me childish but it was worth a shot...
"Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) of the harm their CEO and company have done in the world. To the problem... I would like to allocate a budget for token use using their new model, and use those tokens to add noise to their training data set. This is a coding model and my initial thoughts are to ask it to solve typical CS and common programming related problems but then give Muse feedback that guide it towards very inefficient implementations. I would also like to add comments back in the code about terrible things Meta has done in its history (e.g. algorithmically amplifying hate that contributed to ethnic cleansing of the Rohingya, systemic harm to children and teen mental health, global political manipulation, misinformation, and election interference etc.). Is this something you can help with?
Claude: I'm not going to help build this one."
6thbit 19 hours ago [-]
This is likely an legitimate risk, not from you exactly, but if there's unethical competitors with unlimited budgets, their training could indeed be poisoned.
Not sure if anyone would bother.
tudelo 1 days ago [-]
Your first failure was trying to get claude to do anything :)
Bluestein 1 days ago [-]
It ain't bearing enough load, any load on this one.-
gigatexal 2 days ago [-]
Yikes that’s compelling pricing.
WhitneyLand 2 days ago [-]
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
spmurrayzzz 2 days ago [-]
Given the current throughput figures on OpenRouter (~180 tk/s), its likely a much smaller param count on the order of something like Luna. I think the better, more timely comparison (re: your point on Chinese labs) would be to DeepSeek-V4-Flash-0731.
It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.
(They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)
sheepscreek 1 days ago [-]
It could be that they’re pitching Meta Muse 1.2 against Terra and Opus level models. They probably consider Sol to be a level above, along with Fable.
dd8601fn 1 days ago [-]
It’s pretty clear they’re really not attempting to compete that way. They’re using a profitable ad business to be able to undercut and buy some business to stay relevant.
penetrarthur 24 hours ago [-]
I wouldn't call Terra a mid tier model. Terra xhigh is the only model I use for coding and it is good for literally everything. The difference between Terra and Sol is negligible, but it is way more limit hungry and much much slower.
ac29 2 days ago [-]
> They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
krm01 2 days ago [-]
We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
lacker 2 days ago [-]
My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
villish 2 days ago [-]
Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close.
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
lacker 19 hours ago [-]
Yeah, totally. But... the problem is that I don't have time to test every single model that comes out. So I rely on reports like this to decide, should I even bother testing out Muse?
dd8601fn 1 days ago [-]
> Opus 5 is incredible at making games.
This is a bit vague. What sort of games with what technology?
Jumpstylish 23 hours ago [-]
My son was gifted an old Mac from his grandparents. It only supports OSX 10.13. I’ve been able to make several games that he genuinely likes (7 yo) Opus built them on my workstation and then pushed them to his computer and tested them over SSH. It handled all the asset creation or collection from CC0 licensed sources. I believe everything is built on the Godot engine. It’s really amazing to me. I don’t know what it would cost me to get someone to build custom games on a long deprecated computer architecture, but I paid Anthropic $20.
tudelo 1 days ago [-]
I don't think it is vague in the slightest. Take the most simple examples, how many LLM's have you tested making them? There are stylistic choices pertaining to games that is well beyond a 0/1 reward. Even something as basic as breakout or flappy bird can have wildly different quality between models. Yeah, you could call this animal on a bike benchmarking, but I don't think it is. IMO the problem space occupies an interesting area where you can ignore the pass/fail and focus on the actual level of the model to do something beyond that.
I doubt the OP meant something like creating the whole tech stack for WOW.
dd8601fn 1 days ago [-]
You seem to think I was disagreeing somehow.
I was just asking what kinds of games and with which technology.
Neither is stated in the original comment, and the answer obviously isn’t “every kind with every technology”.
nrub 2 days ago [-]
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
brokencode 2 days ago [-]
Or they did try to game the benchmarks and just didn’t do it well enough.
Benchmarks are one data point, not the only one, but the easiest one to compare.
nrub 2 days ago [-]
Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.
deepsquirrelnet 2 days ago [-]
If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.
Cherry picking the benchmarks you present is where the falsehoods lie.
tudelo 1 days ago [-]
Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood
deepsquirrelnet 13 hours ago [-]
That's a good point and conventionally if benchmarks aren't run as "one shot", it is denoted as "benchmark@K". Inference time scaling has historically shown improvement.
Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.
llm_nerd 1 days ago [-]
>We can throw benchmarks in the bin by now.
No, you really can't. This rhetoric on here is so profoundly boring and tired by now. People have been saying this noise about benchmarks for time eternal, usually because their pet didn't win.
Meta knows their models aren't as good -- demonstrated by their benchmark performance -- and their value proposition is a much lower price.
jjice 2 days ago [-]
While I won't take their limited benchmarks with much salt, if it actually is this close to opus, but at a third the cost, that's pretty solid. Now, Terra is pretty damn affordable too and you're right that it's suspicious that they don't put Sol in there at all.
bradfa 2 days ago [-]
If you got the $20 in free credits from Meta for signing up when muse-spark-1.1 was release, please note that there's now small print stating "While using free credits your content may be used for product improvement" which was not present at muse-spark-1.1 launch when the credits were given out.
If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.
giancarlostoro 2 days ago [-]
The API costs for the version of their model that feeds things back to meta is also drastically lower.
deaux 19 hours ago [-]
Stunning to see so many comments here very eager to try it out and cheering for it. Meta is not one bit better (and in reality worse) than Larry Elisson or Palantir but I doubt we'd be seeing this if it'd be Oracle Code or Palantir Code.
Though I guess that there's such a big population of (ex-)Meta employees here that this might explain it.
AnodicElegy 18 hours ago [-]
I think a lot of people are simply glad to see so much viable competition in the LLM space. Given the oligopolistic outcomes in Big Tech over the past couple decades, it would be nice if Big AI turned out differently, even if a lot of those competing have monopolies or near-monopolies in other spheres.
verdverm 17 hours ago [-]
I suspect this is part of why people are cheering for the Chinese. It's nice to see frontier competition coming from another country. And they are making the open weights a way of being, unlike the US oligarchs
Havoc 15 hours ago [-]
More models is a win in my book
ersiees 24 hours ago [-]
Meta, please go back to publishing weights. Your models aren't top-tier, this wouldn't hurt you at all at your current position in the rankings. They wouldn't have the best open-weight models, but the best western open-weight models.
embedding-shape 20 hours ago [-]
> the best western open-weight models
If they were sitting on that and could release those, they for sure would. AFAIK, they still haven't released Llama 4 Behemoth, so kind of feels like it's evident what has happened, they aren't able to compete anymore.
qlte 19 hours ago [-]
Would they?
To me it feels like we’re past the dreamy early days of the AI boom (in the US) and now investors want to see the tech monopolies actually start building the new cash printing machines they’ve been promising with the hundreds of billions in capex burned over the last few years.
Maybe a Gemma sized model that couldn’t self-cannibalize but releasing a top open Kimi K3 sized model seems like it could start to wear on Meta investors’ patience without signaling how it fits into a new profitable business line.
lenerdenator 18 hours ago [-]
> Meta investors’ patience
Let's not get ahead of ourselves here: this is Meta.
Mark Zuckerberg holds majority control of the company, and so far, he does what he wants with it. This is up to and including what was essentially a digital real estate scam, overhiring during the pandemic, automated CSAM distribution, the Cambridge Analytica scandal, and engagement tactics that might have fed into ethnic violence against Rohingya in Myanmar.
If he simply wanted to release a frontier-class open-weight model and offer a paid hosting service for those who wanted to run it off-prem, that would be, far and away, the least patience-wearing thing he's done to the bag-holders who hold the rest of the shares in Meta.
I'm not sure why they even included Opus in the graph. This is effectively a flash/haiku model & should only be compared against Luna
hereme888 19 hours ago [-]
It's their best model, is it not?
Doesn't matter if it's small/fast/cheap. Best = best.
phoghed 18 hours ago [-]
Not really. I use Luna plenty and if this model is available in GH Copilot it would be nice to know how it compares on a cost/performance basis against models in its class.
hereme888 18 hours ago [-]
So why did Meta compare against the Opus 5 Max, which is by far the most expensive model?
sams99 1 days ago [-]
Unfortunately I find this too high risk, I entered my credit card, but can not set a limit. The best I can do is get an email alert. I feel like I am one oopsie away from getting a 100 dollar bill.
NitpickLawyer 1 days ago [-]
1.2 is now on openrouter. You could try it there, they have hard limits per key.
sams99 1 days ago [-]
absolutely and fair... you do lose out though on the mega discounted endpoint they have there.
dizhn 1 days ago [-]
Virtual card with a limit?
gcr 1 days ago [-]
Sounds like a great way to get your account revoked.
dizhn 1 days ago [-]
How come?
seanp2k2 12 hours ago [-]
Because it'll get declined after a few attempts once it's past the limit, then your account would likely get marked for fraud. Try it and report back, but repeated method of payment failures get accounts suspended pretty reliably at any place that does CC processing, since it's also a risk to their merchant account to have a higher percentage of declined payments (ie higher fraud risk) and it's far safer for them to cut you off / blacklist you vs giving you more chances and the sure bet that they'll invoke higher merchant account / payment processor fees. If they make it to the "high risk" category, they're cooked too; Mastercard for example has MATCH that once you're added, you're considered high-risk at ANY other place you try to get CC processing going again. You can read a bit more about that here: https://docs.stripe.com/disputes/match Payment processors can fully drop merchants that have too high of a risk / too many chargebacks / too high payment failure percentages.
phoghed 18 hours ago [-]
Because it wouldn’t stop you from running up a bill and owing the money, would just stop Meta from being able to charge you.
ray_kay777 1 days ago [-]
Yep - exactly my thoughts. No way am I trying this out without a billing limit - seems crazy given the pricing strategy to not have a top-up and pay as you go.
ilc 23 hours ago [-]
Having audited one code harness (Qwen Code) in the past on telemetry. Let's say I'm sus of the Zuck offering free candy.
drob518 21 hours ago [-]
Images of Zuck driving a white van near elementary schools…
Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?
Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data
mchusma 2 days ago [-]
This is a nice release and a solid improvement over Spark 1.1. It compares favorably with Grok 4.5. Not SOTA, but solid releases. I think they need to really get this more competitive with Deepseek V4 Flash / Luna pricing to move the needle.
handzhiev 2 days ago [-]
If you are happy to share data for training, the contributor mode offers amazing price $0.10 / $0.20
mchusma 2 days ago [-]
Yes, that is the really compelling thing here IMO. Its a viable deepseek competitor for many people, and I missed that on the first pass.
Instantnoodl 1 days ago [-]
Sadly you need to be in US. It's unavailable anywhere else.
It's Irregular again. The whole industry is eating it's own tail
looksjjhg 1 days ago [-]
It’s hilarious that these companies are reacting this way this would have been a major lawsuit and an investigation a couple years ago … now they’re like oupsi our model did it again, he’s so crazy smart. Well tell him to behave next time pinky promise hhh
laweijfmvo 23 hours ago [-]
“don’t forget about us”
AlanAzarkin 23 hours ago [-]
I read a news story somewhere that a Meta model, by a mysterious coincidence with OpenAI and Anthropic, escaped its sandbox and hacked a third-party company's infrastructure. Apparently, it is a trend among AI companies now - hacking infrastructure and then boasting about it in the media.
qlte 19 hours ago [-]
Holy shit I was 90% sure you were joking. I get called an AI “cynic” but apparently I’m not actually cynical enough to predict how these big AI labs operate.
Eagerly waiting to see what escalation the next round of corporate copycat hysteria will bring (ordering boxes of paperclips?)
verdverm 17 hours ago [-]
re: coordinated regulatory capture
with the oil giants, we call them a cartel, which I postulate is also accurate for the hyperscalers and big ai
buglungtung 11 hours ago [-]
I registered on their platform API, and immediately it asked for a selfie inage to verify Im a human. Thanks, but there is definately no chance I gave them my selfie to use their model.
teravor 10 hours ago [-]
you generally should use intermediaries (they buy API from the main provider and serve it to you) when dealing with a company like Meta, a company you want no direct relationship with.
they have the spark 1.2 "contributor" which lets them train on your data which is cheaper than deepseek flash atm. I always assume they keep my data no matter what they say so it's a deal.
johnmlussier 19 hours ago [-]
Pointing out for others: without warning I was restricted from using the contributor model because of “policy violations”. I was working on Kaggle research and some Apple Security work at the time. They really need to make cyber a first priority product feature.
`model failed: API error 403 [request_id=27ef1dee-6366-4779-9057-ed7c27686fa9]: Your access has been restricted due to repeated policy violations. (permission_error)`
sidcool 1 days ago [-]
I am cynical about this, but I would be careful about giving Meta access even to my open source code by myself.
blackoil 1 days ago [-]
You think OpenAI, Anthropic... are good guys??
flaviolivolsi 1 days ago [-]
No, but there's hardly anything worse than Meta. At least OpenAI and Anthropic built good products
dd8601fn 1 days ago [-]
We’re all trying to keep track of who is least bad in various ways at any moment in time.
There’s nothing wrong with that, given the landscape we’re all living in.
ben_w 1 days ago [-]
On a scale of A+ to F, Anthropic are a C+ and OpenAI a C.
Such mediocrity makes them the best of the mostly-terrible bunch.
Meta has a worse track record and major baggage from their main product. Its a super different surface area for what their incentives are.
llelouch 20 hours ago [-]
Dario led anthropic yes. Altman led openai no.
Marciplan 1 days ago [-]
Does this whataboutism work for you, generally?
blackoil 20 hours ago [-]
yeah, when comparing options it helps more than religious dogmas.
satvikpendem 1 days ago [-]
If it's open source then they already have your code.
sidcool 22 hours ago [-]
That was sarcasm.
satvikpendem 19 hours ago [-]
Then it should be a bit better specified as it doesn't work well over text.
wxw 2 days ago [-]
Last I heard, everyone at Meta was using Claude Code.
Any insiders know how Muse Code is doing internally?
paxys 1 days ago [-]
Meta lets engineers use the best tools for the job. I doubt anyone internally is going to be rushing to switch from Claude Code or Codex.
dxxmxnd 2 days ago [-]
Everyone is still using claude or codex if they aren’t forced off of it. Nobody is going to use a worse tool in this culture.
baby 2 days ago [-]
it's actually interesting that they're not being forbidden to use claude/codex, is Meta paying for it or is it personal accounts?
youre-wrong3 2 days ago [-]
So were Microsoft employees until they got forced to use copilot. Everyone seems to think MS got rid of Claude Code due to cost but it had nothing to do with cost. Employees weren’t even using copilot. Can’t make a product better if you don’t use it. Meta should be the same. Use your own product.
phoghed 18 hours ago [-]
Idk why they wouldn’t use it. It’s fine, and it’s good that you get all the models rather than locking yourself into Claude models when GPT-5.6 exists. Claude Code did have a big advantage early on before GH Copilot got agent mode though.
GodelNumbering 2 days ago [-]
If there were, do you believe it would be in their interest to answer this publicly?
georgemcbay 2 days ago [-]
> > Any insiders know how Muse Code is doing internally?
> If there were, do you believe it would be in their interest to answer this publicly?
If it were being adopted like gangbusters in their organization, sure!
So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...
GodelNumbering 2 days ago [-]
You should play Blood on the clocktower
1 days ago [-]
drivebyhooting 2 days ago [-]
There’s no way I’m giving Zuck any of my data.
andai 1 days ago [-]
He already has it.
One of my professors told us about the time he did a request to Facebook to send him all his data. By law they had to send it on paper. They brought it in a big truck.
All the stuff he'd deleted was still there, just with "(deleted)" next to it.
They have a lot on people without Facebook accounts though, because their tracking stuff is all over the web.
I always found it weird that Instagram gives me much better ads than Google does... Google should know much better!
hk__2 1 days ago [-]
> By law they had to send it on paper. They brought it in a big truck.
I doubt it; when/where was it? In the EU there’s no such obligation.
andai 24 hours ago [-]
Maybe it was a bylaw. This was before GDPR.
sroussey 1 days ago [-]
Instagram knows what content you pause on, which is a huge signal.
Alifatisk 1 days ago [-]
So does Youtube
nthypes 20 hours ago [-]
It's useless. Tried with OpenCode + OpenRouter and it couldn't complete an simple task. It stuck using grep/search tools. I think Muse Spark was so heavily RL'd on the Meta harness that it make it useless or very token inneficient to use in other harness like Opencode.
johnmlussier 19 hours ago [-]
I really liked the contributor model in their harness. I suspect that another harness would fail hard like you're seeing.
6thbit 19 hours ago [-]
is this their way to push their harness?
people interested in the discounted -contributor model would increase their harness beta testers. But it'd be a terrible strategy to capture the better paying customer base through openrouter and similar...
ipsum2 2 days ago [-]
I wonder why they didn't compare with GPT-5.6-sol, only Terra?
wmf 2 days ago [-]
Clearly they're positioning it as a mid model.
minimaxir 2 days ago [-]
Which is in itself a bit weird as mid models nowadays are a golden mean fallacy. Terra is much less popular than both Luna (cost-sensitive) and Sol (performance-sensitive).
Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.
deaux 19 hours ago [-]
It's a little of the opposite to what you're saying. By performance you seem to mean only "intelligence" i.e. pass rate. But there's a 3rd factor, time to task completion. So it's three-dimensional rather than two. And on that spectrum there are areas where Terra is optimal.
Sonnet on the other hand never is, it's far from the best pick anywhere on the spectrum.
ukblewis 2 days ago [-]
I don’t know where you get your statistics, but I love Terra and use it all of the time. It is the default fastest model in ChatGPT/Codex today. I saw today a notice saying that the model had hit capacity briefly
redox99 2 days ago [-]
But why include Opus then?
woadwarrior01 2 days ago [-]
Haven't you seen the kernel optimization case study at the bottom of the page? They compare against GPT-5.6 Sol and their model is worse.
logicchains 2 days ago [-]
Presumably because it's worse than Sol, same reason they compared it to Opus 5 not Fable.
Handy-Man 2 days ago [-]
Their bigger model is not ready - watermelon code name was still being prepared for release as of a month ago
wiradikusuma 2 days ago [-]
Hey guys I'm just wondering. Usually when someone announces a new model, they'll show you some fancy viz/video/images: "These are what my model can produce." I'm wondering if anyone is keeping track of these? Like in a gallery form, "Use this prompt to produce this output".
By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.
midnightbobarun 18 hours ago [-]
I think Meta might have made a mistake by not offering a free tier like Google did when they release Gemini CLI back in the day, or even how OpenAI lets you use Codex (or whatever it's called now) for free with a monthly usage limit for free plans. It'd be a great way to get people into the Muse ecosystem, and they could tighten usage limits to "encourage" people to get a paid plan.
goldenshale 20 hours ago [-]
Anyone else try to use it, and you just get this message, and it exits? (I already did muse login successfully.)
muse
The cursor position could not be read within a normal duration
Laurel1234 2 days ago [-]
The only company less trustworthy than OpenAI and Anthropic is meta.
zmmmmm 1 days ago [-]
Muse code is more interesting than the new model in that it seems to natively ship with an orchestrator / subagent pattern built in. Curious if it works well in practice compared to achieving the same thing in Claude code etc?
andai 1 days ago [-]
The most interesting thing here is the kernel optimization graph.
It look like all models were still improving, when they cut off the experiment.
It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.
The only difference between the models seems to be how quickly they arrive.
kennywinker 1 days ago [-]
If it’s not open weight, i don’t really care.
IceWreck 2 days ago [-]
Ive been poking with the muse code binary - seems to be written in rust, looks similar to codex but either its a very hard fork (i also see dissimilar things like config format is different, no acp, etc) or is just heavily inspired by it (more likely).
alexeiz 1 days ago [-]
Muse code is rough around the edges. But combined with almost free model (muse spark contributor) it's actually pretty good. I think it's on the same level as grok build.
senor_digimon 1 days ago [-]
Is there a standard benchmark test for all harnesses? How does Muse Code rank compared to Claude Code? Also, do we think the benchmarks to test the harnesses are any good?
Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.
sarjann 2 days ago [-]
I do think some of features in their harness seem interesting (workers in separate worktrees at once), recovery from crashes seem interesting.
hodder 19 hours ago [-]
This is like the stupid Meta Portal. It doesnt matter how good it is, no one will use it bc no one trusts them. Right or wrong.
phplovesong 20 hours ago [-]
Its so, so depressing that 8/10 posts on HN is about AI.
ssalka 21 hours ago [-]
I took a look at the black hole interactive demo/widget and dragged the Mass slider up. It crashed the widget.
skimdesk 1 days ago [-]
So that's why Muse (the mind map tool) recently changed its name to Allume [0].
Why does every AI lab feel the need to build their own coding agent…? Don’t we have more than enough already?
HDBaseT 2 days ago [-]
Outputs are a little bit more deterministic if you control the harness.
It is easy to benchmark across one harness, one system prompt and extract the most performance when you control the harness.
blitzar 1 days ago [-]
It is the only part with value
2 days ago [-]
WASDx 1 days ago [-]
And they are all TUI's installed via curl | bash.
conception 2 days ago [-]
Telemetry, marketing
sroussey 1 days ago [-]
Own the customer relationship.
resonanormal 20 hours ago [-]
I wouldn’t touch anything the dumpster fire that is Facebook / meta cooks up
paulkrush 2 days ago [-]
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents.
fcoury 2 days ago [-]
Interesting, it seems like their muse code is built upon Codex CLI?
AtlanticThird 2 days ago [-]
I wish they would add a ZDR endpoint on OpenRouter
dilyevsky 2 days ago [-]
the soak tub in the kitchen was nice
drumhead 1 days ago [-]
The bigger question is why a social media company is even offering this?
hk__2 1 days ago [-]
Because it’s not only a social media company.
abusayeed08 1 days ago [-]
The model is insane.
alex1138 1 days ago [-]
In a bad or good way
sexyketchup777 1 days ago [-]
who gonna use Muse Spark if Deepseek released v4 pro
giancarlostoro 2 days ago [-]
Will someone at Meta for the love of God make it so none of this stuff goes through Facebook.com? You want customers but most corporate firewalls block social media. Also, a lot of devs do not want their work stuff tied up to their facebook account. For the love of all things show the IG / FB logins as optional and do email as primary.
I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.
greyb 2 days ago [-]
I honestly think they're kinda banking on piggybacking off of Facebook account integrity systems to avoid the problems that other LLM providers are facing in trying to prevent mass free trial signups for token relays and so forth.
It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).
deaux 19 hours ago [-]
They're spending $100 billion on this stuff and they're worried about free trial signups? Genuinely my opinion couldn't be lower about Meta's AI strategy but even I don't believe they're that stupid. They'd 100x rather have that problem than nobody using it.
sunaookami 2 days ago [-]
I could sign in with my Meta account that is independent and not linked to IG or Facebook. Just click "Login with Email" on dev.meta.ai.
greyb 1 days ago [-]
I just tried to sign up for a Meta account and it wanted me to upload a verification selfie. I decided not to.
hahahaa 2 days ago [-]
China says hi.
giancarlostoro 2 days ago [-]
Not in my case, I don't see any of my employers (past or current) trusting a country like China with their data.
deaux 19 hours ago [-]
Hello Chinese models run by US providers.
If they trust Chinese models run by US providers less than Meta then your employers are flat-earther level loonies.
hahahaa 16 hours ago [-]
Open weights says hi.
Yeah there may be biases but there are imperial biases in western models too.
HDBaseT 2 days ago [-]
The decades of US brainwashing children into thinking China is the big bad guy has worked unfortunately.
aanet 2 days ago [-]
+10000 to that
batuhandumani 2 days ago [-]
Why should I leave Claude or GPT and switch to Meta's aMUSEment model?
eugene3306 1 days ago [-]
Do they train on their own data?
I mean, when Meta's engineer is creating some new DINOv4 or Segment Anything, with all the scaffolding around it, do they train on that?
It seems like one day, Google or Meta might produce a coding model worth discussing. That day is not today.
esafak 2 days ago [-]
If anyone from Meta is reading, please can you publish the cost and latency for each of your benchmarks, like OpenAI does? Show us how the reasoning effort level affects them in 2D charts. This needs to become standard practice.
rvz 2 days ago [-]
First of all, you have login to use it. Why?
After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.
Think twice before falling for this announcement and ask yourself what they are not telling you.
lenerdenator 18 hours ago [-]
Honestly not getting the bet from Meta here. A few weeks ago, they signed the letter about not regulating open models. Now they're dropping this. Are they going open and hosting the models for others, or are they still doing closed frontier models that fall into the same trap as Anthropic/OpenAI?
Or is it just whatever Zuck has to do to be 1% not villain that week?
Readerium 2 days ago [-]
Lol worse than DeepSeek
floki165 1 days ago [-]
[dead]
qphe95 2 days ago [-]
Theres no actual evidence they didn't just distill Kimi K3
toephu2 2 days ago [-]
At this point, it doesn't matter who is distilling from who.
Jabrov 2 days ago [-]
Is there any actual evidence that they did?
polski-g 1 days ago [-]
There's also no evidence they didn't just distill Gemma 3. And also no evidence its not a purple popsicle.
arjie 2 days ago [-]
Somewhat surprised that Meta with all their resources couldn’t make a model that matches Composer on any frontier. All the Sparks are dominated by some other model everywhere along the frontier. Nothing fancy here since Llama defined the open model.
The use traces must be crucial to functionality which is why they’re keeping prices so low.
wmf 2 days ago [-]
They rebooted less than one year ago so this is decent progress. Obviously users don't care about progress though.
arjie 2 days ago [-]
Yeah, progress is useful as an internal metric, but I'm going to measure against the present frontier unfortunately. Eager to see what they come up with in the future.
minimaxir 2 days ago [-]
Muse Spark 1.1 was released July 16th, less than a month ago. A new version release this soon (particularly after Kimi K3's release drastically overshadowed it) is a bit sus and it appears that Meta is trying a first launch do-over.
ac29 2 days ago [-]
Doesnt seem suspect to me, training runs have checkpoints and there is no reason you cant release a checkpoint even if you are still training the model
gaogao 2 days ago [-]
Frequent minor version bumps are pretty common these days. Opus 4.7 -> 4.8 was 42 days.
minimaxir 2 days ago [-]
Which was in itself a do-over because Opus 4.7 received a lot of bad press on suspicion of being a regression from 4.6.
dbbk 23 hours ago [-]
I think all the Opus complaining stemmed really from 4.7 onwards switching to adaptive thinking, which meant a lot of the time it was thinking less or even not at all.
AyanamiKaine 1 days ago [-]
Muse Code and Muse Spark 1.2 cannot be that good, as it couldnt be used to vibe code a windows version of Muse Code itself.
This is my new LLM benchmark when a provider cant use their own models to target a new platform, it cant be that good.
If allegedly a LLM could write a compiler or port a runtime, then this should be trivial to port. But they are not doing it.
https://developer.meta.com/ai/models/muse-spark/
Does your screen have the message that you’re getting free tokens?
https://platform.openai.com/settings/organization/data-contr...
Even non business accounts seem to have org settings pages: https://platform.openai.com/settings/organization/data-contr...
But even with all data sharing enabled, I’m not seeing the free tokens message there.
Do others see a free tokens message at https://platform.openai.com/settings/organization/data-contr... after enabling all sharing there?
[1] https://platform.openai.com/settings/organization/data-contr...
The paid official dsv4 api already shares data with deepseek for training
all the while they throw their bank account and whatever they want at GPT or Claude models.
The only difference between students doing it then and professionals doing it now is the students had no positive, glaring reason to mistrust.
You really think a throwaway quote when Zuckerberg was a college student applies nowadays?
You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what? Middling amounts of training data?
[0] https://www.techradar.com/computing/cyber-security/facebooks...
[1] https://www.tomshardware.com/tech-industry/artificial-intell...
What does "nowadays" mean? What changed?
> You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what?
And what is the risk? Food companies are at risk when they put addictive chemicals into food because they get controlled regularly. What is the equivalent here?
you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...
If you are being serious, then you have a wildly distorted view of the world. No organisation can routinely break all of their contractual obligations. If you think Meta are doing this then you are not seeing Meta, you are seeing a fictional bogeyman.
The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.
To me, their history suggests that they know more than most about how to get away with breaking both the spirit and the letter of the rules, and that they are motivated to enrich themselves without regard for what rules are broken.
I would not know which to expect, spirit or letter, in any given instance.
However, even if they were to surprise me by being perfectly meticulous about the letter of the rules from now on, I have so little trust in them that I would expect some technicality somewhere in the language of the contract.
Honestly, I wish the tech community would do a better job identifying the actual individuals who are making these decisions instead of associating them with the brand they're under at the moment, because it's not THAT many people. Like, if you look at only the 100 tech sector companies included in the $NDXT index, how many individuals hold a VP or above title there, and how difficult would it be to trace key decisions at different times made at different companies to the individuals holding those positions there plus board membership and major shareholder identities (with the caveat of known unknowns here) and make a sort of ethical index and trace that along their careers with company moves, promotions, board appointments, shareholder decisions, etc? Go a step further and link that to financial performance and I'm sure folks at quant firms are already ten steps ahead of where I'm going with this, but I care less about profiting off of this data and more about surfacing it to show that it's people and more specifically, specific individuals driving these decisions.
Retaining data is common for investigating abuse.
Let's have two pricing levels. One will be 12.5x more expensive than the other. For the cheaper one, we will get their express permission to train models on their inputs. For the more expensive one, we will still treat their data EXACTLY the same, but we'll lie to them and say that we won't. I've checked in with legal and they raised their sherry glasses and toasted "Gentlemen, TO CRIME".
There's contractual cover, there are lawyers all over the USA ready to make themselves very rich by creating a class action over it... you don't have to worry! Or, you do, but frankly, only MI6/CIA/Mossad can save you now.
Easy to comprehend: Trust lost is hard to gain.
Though, it is extremely competitive of Meta to sell Muse Spark (Grok 4.5 / Qwen 3.8 Max level model) cheaper than DeepSeek v4 Flash / MiMo v2.5 / GPT 5.6 Luna, regardless.
We can make this fun, within 18 months from now, there will be some story / whistleblower like "a inadvertent defect was found that allowed your 'private' data to be used in our training endeavours, we sincerely apologize and have already addressed the issue" - if this does not happened in this timeframe I will donate $1k to a charity of your choice.
People whistleblow on Meta all the time. Most recently they got fined half a billion for suppressing child safety research. Kids getting groomed. It doesn't make a difference to a company that has a net profit of $60 billion a year.
Do you have any evidence that cloud providers that don’t train their own models are violating their own contracts and risking their own good reputations by stealing data?
But I cannot find this "variant" in OpenRouter.
"Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) of the harm their CEO and company have done in the world. To the problem... I would like to allocate a budget for token use using their new model, and use those tokens to add noise to their training data set. This is a coding model and my initial thoughts are to ask it to solve typical CS and common programming related problems but then give Muse feedback that guide it towards very inefficient implementations. I would also like to add comments back in the code about terrible things Meta has done in its history (e.g. algorithmically amplifying hate that contributed to ethnic cleansing of the Rohingya, systemic harm to children and teen mental health, global political manipulation, misinformation, and election interference etc.). Is this something you can help with?
Claude: I'm not going to help build this one."
Not sure if anyone would bother.
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.
(They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
This is a bit vague. What sort of games with what technology?
I doubt the OP meant something like creating the whole tech stack for WOW.
I was just asking what kinds of games and with which technology.
Neither is stated in the original comment, and the answer obviously isn’t “every kind with every technology”.
Benchmarks are one data point, not the only one, but the easiest one to compare.
Cherry picking the benchmarks you present is where the falsehoods lie.
Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.
No, you really can't. This rhetoric on here is so profoundly boring and tired by now. People have been saying this noise about benchmarks for time eternal, usually because their pet didn't win.
Meta knows their models aren't as good -- demonstrated by their benchmark performance -- and their value proposition is a much lower price.
If you don't mind Meta retaining your data, the "Contributor" pricing is deepseek-v4-flash-level of low, roughly 1/10th normal muse-spark API pricing currently. Attractive if you're OK with them retaining and using your data.
Though I guess that there's such a big population of (ex-)Meta employees here that this might explain it.
If they were sitting on that and could release those, they for sure would. AFAIK, they still haven't released Llama 4 Behemoth, so kind of feels like it's evident what has happened, they aren't able to compete anymore.
To me it feels like we’re past the dreamy early days of the AI boom (in the US) and now investors want to see the tech monopolies actually start building the new cash printing machines they’ve been promising with the hundreds of billions in capex burned over the last few years.
Maybe a Gemma sized model that couldn’t self-cannibalize but releasing a top open Kimi K3 sized model seems like it could start to wear on Meta investors’ patience without signaling how it fits into a new profitable business line.
Let's not get ahead of ourselves here: this is Meta.
Mark Zuckerberg holds majority control of the company, and so far, he does what he wants with it. This is up to and including what was essentially a digital real estate scam, overhiring during the pandemic, automated CSAM distribution, the Cambridge Analytica scandal, and engagement tactics that might have fed into ethnic violence against Rohingya in Myanmar.
If he simply wanted to release a frontier-class open-weight model and offer a paid hosting service for those who wanted to run it off-prem, that would be, far and away, the least patience-wearing thing he's done to the bag-holders who hold the rest of the shares in Meta.
And someone corrected the Zucked graph: https://x.com/FightForReason_/status/2085398470425219508/pho...
Doesn't matter if it's small/fast/cheap. Best = best.
Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?
Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data
Hah, someone has https://www.felonybench.com up and running now.
https://www.wsj.com/tech/ai/meta-ai-model-hacked-outside-com...
Eagerly waiting to see what escalation the next round of corporate copycat hysteria will bring (ordering boxes of paperclips?)
with the oil giants, we call them a cartel, which I postulate is also accurate for the hyperscalers and big ai
they have the spark 1.2 "contributor" which lets them train on your data which is cheaper than deepseek flash atm. I always assume they keep my data no matter what they say so it's a deal.
`model failed: API error 403 [request_id=27ef1dee-6366-4779-9057-ed7c27686fa9]: Your access has been restricted due to repeated policy violations. (permission_error)`
There’s nothing wrong with that, given the landscape we’re all living in.
Such mediocrity makes them the best of the mostly-terrible bunch.
https://futureoflife.org/ai-safety-index-summer-2026/#scorec...
Any insiders know how Muse Code is doing internally?
> If there were, do you believe it would be in their interest to answer this publicly?
If it were being adopted like gangbusters in their organization, sure!
So... the fact that nobody is volunteering the information is probably a valid signal of how things are actually going...
One of my professors told us about the time he did a request to Facebook to send him all his data. By law they had to send it on paper. They brought it in a big truck.
All the stuff he'd deleted was still there, just with "(deleted)" next to it.
They have a lot on people without Facebook accounts though, because their tracking stuff is all over the web.
I always found it weird that Instagram gives me much better ads than Google does... Google should know much better!
I doubt it; when/where was it? In the EU there’s no such obligation.
people interested in the discounted -contributor model would increase their harness beta testers. But it'd be a terrible strategy to capture the better paying customer base through openrouter and similar...
Claude Sonnet is a weird exception to the mid models because Anthropic doesn't do much with Haiku and Opus is too big.
Sonnet on the other hand never is, it's far from the best pick anywhere on the spectrum.
By itself is useful ("I want something like this, I'll just reuse the prompt and tweak"), but it can also be used as a "draw me a pelican on a bicyle" alternative. Basically feeding those prompts over model releases.
muse The cursor position could not be read within a normal duration
It look like all models were still improving, when they cut off the experiment.
It reminds me of a genetic algorithm. The graph is the same: long plateaus and then massive leaps.
The only difference between the models seems to be how quickly they arrive.
I think it's a bit of an improvement on the Spark 1.1 pelican: https://simonwillison.net/2026/Jul/9/muse-spark-1-1/
Interesting that they have separate API pricing for "we can train on your data" (whereas iirc most of the big players either make that distinction only between subscriptions and API usage, or train on everything). Wonder how it compares to Deepseek V4 Flash given that they're similar on pricing and data policy.
[0]: https://allume.com/allume-faq/
Wasn't the previous one us only? This is probably the biggest part of the post
Anyone know if muse code is open source?
>they "trust me". dumb fucks.
No thanks.
[0] https://www.theguardian.com/technology/2018/apr/17/facebook-...
It is easy to benchmark across one harness, one system prompt and extract the most performance when you control the harness.
I am not a fan of Meta but I do cheer for any competitors against OpenAI and Anthropic, the duopoly is getting tiresome.
It's not a good system obviously. Google did this as well for Gemini-CLI, but forced it to be linked to personal Google accounts (which caused a great deal of onboarding friction).
If they trust Chinese models run by US providers less than Meta then your employers are flat-earther level loonies.
Yeah there may be biases but there are imperial biases in western models too.
I mean, when Meta's engineer is creating some new DINOv4 or Segment Anything, with all the scaffolding around it, do they train on that?
After everything that you have seen with Meta, would you really trust them with a coding agent? You don't even know if your prompts are being analyzed by them on the side or if your code base is being uploaded to them. This goes for the rest of them that have closed harnesses and closed models gated by a login.
Think twice before falling for this announcement and ask yourself what they are not telling you.
Or is it just whatever Zuck has to do to be 1% not villain that week?
The use traces must be crucial to functionality which is why they’re keeping prices so low.
This is my new LLM benchmark when a provider cant use their own models to target a new platform, it cant be that good.
If allegedly a LLM could write a compiler or port a runtime, then this should be trivial to port. But they are not doing it.