Hacker Newsnew | past | comments | ask | show | jobs | submit | larodi's commentslogin

> you can forget about anything coming of this.

first of all - a lot of people are definitely not forgetting it. perhaps many more are waking up to the fact. when so many people wake up to the fact that a massive theft of intellectual property IS what enabled present day AI, they will inevitably refuse to a) publish that much openly; b) respect any kind of copyright claims imposed by those who perpetuated, facilitated, enabled the theft.

so, really, a lot will be coming out of it, we like it or not.


The whole original conversation was very smelly from the very start, 20k stars included on the GitHub page with lost history.

The Bend1 launch was very successful, with 1000+ hn upvotes [0] and with several youtube videos with up to 1M+ views [1] [2] - that's where a lot of the popularity came from, I believe. They just recycled the Bend1 repo for Bend2 even though it's a completely different language.

[0] https://news.ycombinator.com/item?id=40390287

[1] https://www.youtube.com/watch?v=HCOQmKTFzYY

[2] https://www.youtube.com/watch?v=NaytZOiX3fs


> ...on the GitHub page with lost history.

> They just recycled the Bend1 repo for Bend2 even though it's a completely different language.

Why would you kill the source history of Bend1 completely if it has 20K stars? You can only do that if Bend1 has no users at all right?

Does that meant that Bend1 isn't actually successful in its own right but rather only as a marketing project?

I wasn't one of the suspicious people but I am now.


(author here) Bend1 indeed has no significant active userbase

I don't think that means it was "unsuccessful" in the sense you imply, though, because the project was never meant to be used in production. It was there to display a milestone (running inets on the GPU) and I was very clear it wasn't ready to be used yet. For example, it had only 24-bit integers, a 2 GB memory cap, and other limitations that made it unpractical. I still don't know why it has so many stars. I posted it to hacker news and that just happened. I guess it just went viral without really being ready yet, which got us to where we are now.

Anyway the commit history is back now. I apologize for nuking it


[flagged]


It _is_ sus, but you should be willing to forgive. It's not like it was done out of animus, and I assume that an academic/foreigner is unaccustomed to how production software conventions work. Their baseline is less important, more important is how quickly they learn/adopt "best practices". Y-intercept, slope, etc.

I say this as literally the first HN comment to call them out on this.


it becomes more sus as the superfans come out

the author here is providing commentary on the vibe coding era, bend is just one example of people having ai build things for them they don't understand or haven't researched sufficiently, at least start with a vibe-search skill before the vibe-coding begins

Ai is only good at a task when the human driver is good at that task, and this is far more narrow than most realize. Someone who knows how to program will still fail on many programming tasks with agents, the field has far too much for any one person to know.


It sounds even worse. Actually far worse. It's a project already in use and the author just nuked the whole git history?

It's a research language, of the more experimental and niche type.

No one uses it seriously, and the author since then reverted the nuking of the git history.


Does anyone know where they bought the popularity and contributors from? I would like to do this for my joke language to fool unsuspecting users into using it seriously

Probably at the same market where they sell snark. You may want to sell your excess of it, especially that as mentioned there were a popular (as in got to HN frontpage multiple times) version 1 and it's all legit and honest work, getting uncalled criticism?

Please do not shill your crypto here, nobody wants to buy snarkcoin

The history is back...

Yeah 100% those are purchased/fake

Lean4 itself has 9k


I don’t they are purchased or fake.

I can unequivocally claim that Victor receives at least two orders of magnitude more social media engagement, while Lean likely has 4-5 more magnitudes of actual users.

Serious users of Lean simply have no need or reason to star the repository.


It is super amazing that 3 years later, none of the models' weights developed by Anthropic or/and OpenAI have leaked so far. Not a single one.

Windows internal builds have leaked for years, early game versions, GTA videos, secret documents, whatnot. But somehow even though all the whistleblowing, not a single model was leaked. What level of security do these companies have? Do they bring encrypted DVDs to AWS to run the services or really...how's it even possible?


One trivial reason might be the size of the artefacts / hardware requirements? Kimi K3 is ≈ 1.5 TB and requires multi million dollar hardware to run. Compared to e.g game development, I'm guessing that it's not like a bunch of people at Anthropic/OpenAI have the models running "locally".

It's easier to protect a power substation from being stolen then a Rolex watch


Well this concludes then that it’s like a handful of actual engineers and ML ppl that have access to it and have taken all precautions to keep it locked.

Again - many people have so far left these companies and none brought an usb drive out with what very likely does not constitute copyrightable materials in the first place.


Or the ones doing the stealing are so competent (or embedded) we don't hear about it

> Kimi K3 is ≈ 1.5 TB and requires multi million dollar hardware to run.

Not that it defeats your point, but an 8x MI355X node is $350k-400k. The only reason you're paying that much is for the VRAM, too. You could run it with much less compute than what you get in a single card.


That's "only" 11h of download at 300Mbit/s

How long would it be on 56k? I recall having to reconnect to my ISP every three hours to resume downloading an iso back in those days.

SSO and hardware sec keys. And the models are located in very few places. Few if any people have direct access to them. Then due to the size of the models you can detect and stop a theft just by monitoring the egress traffic.

> monitoring the egress traffic

Oh so you mean the thing HuggingFace wasn't doing at all while also allowing any user's arbitrary programs to call out to the open web from prod?


indeed, how does this add up the fact that 20k agents drilling another corpo's headquarters may actually register on a radar. and it seems they did on multiple occasions...?

probably a bit harder to steal terabytes of data, and the weights aren't what people are after anyway - distillation is basically "stealing" a model and you can do it from outside

Publicly…

Windows internal builds and video games are distributed to engineers and testers to run on their local workstations/consoles. Model weights are not.

People working at OpenAI have stock options. People working at MS and Rockstar do not.

Leaking negatively affects investment while the “whistleblowers” are largely just saying “our tech is too good” which increases investment into those companies.

Ultimately, it always comes down to money.


MS absolutely has a couple of stock-based incentives.

is it not visa free already for most of the states anyway?

Generally it is auto approved for Canadians, but you will soon need to do an "ETIAS travel authorization".

Canada actually implemented their ETA way before the EU did theirs.

Yes but with that you’d be able to have a work permit as well!

you'll have to prove equivalence through some Lean4 code perhaps? or some weird clause tree comparisons... good question indeed.

"is this the real thing or is just fantasy"

So proud to see the prof. who oversaw my masters thesis has their paper (and solutions) featured here. :D

Given (A) :

> He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models. They famously ingested plenty of copyrighted material without the permission of those intellectual property holders.

And many people's shared opinion (B):

>> I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.

...

It is very difficult to actually say NO to the fact that (A) was done, which then leads logically to conclusions as (B). But also we should remember that if these two hold (and (A) is an axiom more or less now), then it comes as no surprise that then also all opensource licensing is immediately rendered void and null, as keeping it would contradict (A) and would go against the very common and consequential logic in (B).

Copyright is so dead. And it was not me killing it with a cynical post on HN. Dunno why so many people still fail to face it. There is no way it can exist in its current form, because then immediately (A) happens and (B) follows.


well the OS kernel does quite a lot and stays autonomous for even more hours and seems super smart to me - I don't even know what tis doing most of the time, but it does it very well. autonomy and predictability are different things then. the fact something is unpredictable does not make it more autonomous than something that we can trace the logic for.

The kernel is programmed, and it only ever does exactly what it was programmed to do. So all of its behavior must be specified, and any problem it “solves” was solved by the programmer who told it exactly how to solve that problem. That is not autonomous.

Autonomous is when I tell codex “here is the chat where the user reported the issue. Solve the problem they reported” and I come back later and the problem is solved.


I think the argument rather is that there is not such thing as AI that is not human-alike. The traits that AI exposes, if I get the point made here correctly, are traits that existed in the training material, even though not evident for the human eye. It is a mirror-like reflection of humanity's traits found in the source material, an Egregore* of a sort.

The intelligence emerging, or whatever is emerging, more or less follows human traits, the CoT exposed is very human-like etc. Even though we should not anthropomorphize AI as being sentient or conscious or having a soul, etc., we have to remember that the source material is still not alien writing, or concepts that came from Sirius or emerged from a wormhole.

The concern is really nobody reads (or have ever read) 77 TBs of information in order to summarize the most evident traits in it, and there is a fair chance, that where humanity stands today, there are so many bad/damaging traits recorded in this corpus. Then there is a chance the most important features are on a level (or have dimensionality) that is incomprehensible, that humans are unable to reason about (or only few can), because such amount of information is very hard to reason about, or because we don't usually facilitate fractal dimensions when we reason about text and everyday concepts,... or something even more bizarre.

* https://en.wikipedia.org/wiki/Egregore * https://en.wikipedia.org/wiki/Fractal_dimension


The argument became false as soon as we started reinforcement learning. Simple token-completion is arguably just simulating human behavior. Once you start training it on specific tasks you're training a level of single-minded goal-seeking unlike anything found in biological life.

> a level of single-minded goal-seeking unlike anything found in biological life.

hmm sounds like exact description of viruses, which, although not really alive, are part of/influence biological life for both good and bad. in fact humanity wouldn't be where it is if not for certain viruses.

it can be argued there is a lot of bacteria that is super goal seeking, and perhaps also very narrow minded, if minded at all. and this argument, the single-mided one, goes very far up the ladder of otherwise complex organisms.

unless we all are creationists, and God forgive me for expecting the whole game to be much more complex than just snapping humanity into existence, we can then say that trial and error is what drives evolution, and the idea of preserving life and procreation is veeery single-minded goal on its own.


>the idea of preserving life and procreation is veeery single-minded goal on its own

Prevalence of sub-replacement fertility strongly argues against this.


As far as I know, none of these state of the art AIs have been trained on any specific task. It's as general as possible.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: