Rendered at 20:11:22 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
JPLeRouzic 12 hours ago [-]
I have read somewhere that Transformer architecture has a quadratic cost [0] (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets).
For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:
Apparently none of the people complaining about rust use read past the title because it's the 2nd section of the readme and impossible to miss.
kasumispencer2 13 hours ago [-]
That part is newly added after people complained.
Edit: the claimed reason of "updates" also seems to not exist in code. If the author is going to use LLM for this, the very least they can do is to ask it to check properly before publishing.
hashar 13 hours ago [-]
To quote the readme:
> The implementation language is a detail, and a Python port is welcome.
I guess the post title could have dropped "in rust"
mllev15 16 hours ago [-]
You can just tell when the idea itself was generated
janalsncm 15 hours ago [-]
OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.
You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.
Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?
In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.
yjftsjthsd-h 14 hours ago [-]
I dunno, I could probably be convinced to try a new tool purely on the basis of not having to deal with installing pytorch
intoXbox 13 hours ago [-]
I’m curious, what’s the criticism for PyTorch?
jeroenhd 12 hours ago [-]
I don't think it's caused by PyTorch on its own, but every AI-related Python project I try out locally manages to depend on a version of PyTorch that isn't in my disk cache yet. Having to download a gigabyte of dependencies for every project gets tiresome.
The Rust compile cycle will probably generate a gigabyte of files locally as well, but at least they can be `rm`'d out of `target/` once it's done.
It should be said that for this project that's entirely irrelevant of course, but seeing PyTorch has made me skip over projects on the HN homepage before and probably will again in the future.
tomtom1337 13 hours ago [-]
One criticism is that you have to install the same package, torch, but from different Python indexes in order to install the cpu version or gpu version, on Linux. On windows, `pip install torch` gets you the cpu version. On linux, that gets you a ton of Nvidia extras that take a lot of space.
GPU support should really be a optional extra eg `torch[gpu]` or `torch[nvidia]`.
coredog64 5 hours ago [-]
I would argue that we need a `torch[stubs]` package (or similar) that doesn't install any form of libtorch. The CPU version of libtorch is 160MB compressed and 400MB extracted.
pseudocomposer 8 hours ago [-]
The entire Python ecosystem is horrible and shouldn’t have been as falsely boosted by institutions as it was in the 2010s.
Yes, it got less bad with 3.8 or whatever version added type annotations. But making so much of ML depend on Python has made it distasteful to a lot of devs who would otherwise have contributed more to it.
We really need to move all AI/ML research off PyTorch to Candle or… just anything that isn’t Python or another old-gen, broken language like it.
IshKebab 13 hours ago [-]
Yeah likewise. Pytorch needs to die.
That said I don't know what's wrong with using a Rust AI framework like Candle.
alightsoul 15 hours ago [-]
Apparently op cares a lot about speed, which is fine, but ML researchers care about correctness first, speed second. And it makes sense, because they are not as resource constrained as OP.
janalsncm 14 hours ago [-]
Most PyTorch tensor operations are cython not python. So imo rewriting in rust is not going to have an enormous speed up. If that really was the concern we should see a throughput comparison vs PyTorch or something.
skeledrew 12 hours ago [-]
Are you trying to make an argument here that speed is more important than correctness? I'm finding it difficult to interpret - the purpose of - this comment otherwise, and if you are, I'd consider such an argument pretty wild.
lunchbucket 12 hours ago [-]
They aren't using Rust for speed.
nicman23 14 hours ago [-]
yeah that is why they compute in fp32 lol
a19486 15 hours ago [-]
Homie just had some tokens to burn at the end of the month and “in Rust” is pure HN clickbait.
Marcuss2 15 hours ago [-]
How would it compare to current state of the art tested state space layers like Kimi Delta Attention?
saglogog 14 hours ago [-]
Nice idea actually, I always wondered why there was no actual programming in ML!
skeledrew 12 hours ago [-]
Well, to be a bit pedantic, it's called Machine Learning, not Machine Programming. But also learning is just the other side of the skills transfer coin (that teaching, or programming, is on), so it's more a matter of perspective. And then, actually there has always been - traditional - programming in ML: the data has to be prepped, architecture created, etc.
meredithbloom 13 hours ago [-]
> The architecture is the claim here.
Weird turn of phrase very typical of AI.
kasumispencer2 15 hours ago [-]
Isn't this just RNN and nearly nothing about this is actually new?
anon291 16 hours ago [-]
Seems similar to a neural turing machine.
sparticle62 15 hours ago [-]
That's the closest prior work, yes, and the memory half sits squarely in that lineage. Three differences worth naming. Reads are content-addressed in the Poincare ball rather than by cosine similarity in R^n, which is what lets a bounded top-4 read cover a hierarchy of contexts instead of a flat neighborhood. There is no controller emitting explicit read and write heads; writes are novelty-gated with a refractory counter that rate-limits overwriting the same slot, so repeated contradictory updates do less damage. And the fast weights do not stay external forever, they get consolidated into the recurrent transition matrix by closed-form ridge regression.
The backbone is also a selective SSM rather than an LSTM controller, so the sequence half is closer to Mamba than to an NTM. DNC comparisons are fair too, and I should have cited both in the readme.
throwaway_7274 6 hours ago [-]
Is this performance art? Pretty disrespectful to meat-proxy a response to someone engaging with your, uh, work.
Edit: plausibly there’s no human in loop here at all, in which case fair play, I guess.
purple-leafy 16 hours ago [-]
“In Rust” man of all the cliche title baits, I hate this one the most.
Why did you choose Rust? Why does that matter?
Nothing wrong with Rust. Lots wrong with the bandwagon that “in Rust” somehow adds value.
janalsncm 15 hours ago [-]
Rust has its place but python is default for this kind of thing. By doing it in rust, they have now changed two things: the implementation of the transformer, and this new model.
echelon 16 hours ago [-]
Counterpoint: I only clicked on it because it said "in Rust".
IshKebab 13 hours ago [-]
Rust matters because it means you can actually deploy it without going insane. I wonder if the anti-Rust zealots have actually ever used pip, especially on Windows.
He could have said "... not written in Python" - would that have been acceptable?
No he could have left out languages entirely. It’s not relevant for actual discussion.
A non transformer language model written from scratch is more interesting than
“I did X in Python/Rust/Flavour-of-the-month-thing”
IshKebab 6 hours ago [-]
I would say the ability to actually use software is kind of relevant to said software. There's definitely software I wanted to use but gave up on because it was just impossible to actually install. Hell I just spent half of today coaxing Claude to fix some Kotlin software so that it worked with the version of Java I have. Let me tell you that would have taken a good few weeks without AI, even if I was unfortunate enough to be an expert in Maven and Gradle. "Written in Rust" would have saved me a huge amount of (pre-AI) time for that software.
anon291 16 hours ago [-]
Yeah I read the title and did a double-take. It's an uninteresting choice for systems like these. The more pressing concerns, which go undescribed, are the exact mathematical choices behind the actual model. Rust provides almost zero value here because tensor stuff is all just 2d-arrays of floats for the most part.
For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:
[0] https://aclanthology.org/2023.findings-emnlp.936/
https://arxiv.org/abs/2503.00392
https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...
Edit: the claimed reason of "updates" also seems to not exist in code. If the author is going to use LLM for this, the very least they can do is to ask it to check properly before publishing.
> The implementation language is a detail, and a Python port is welcome.
I guess the post title could have dropped "in rust"
You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.
Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?
In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.
The Rust compile cycle will probably generate a gigabyte of files locally as well, but at least they can be `rm`'d out of `target/` once it's done.
It should be said that for this project that's entirely irrelevant of course, but seeing PyTorch has made me skip over projects on the HN homepage before and probably will again in the future.
GPU support should really be a optional extra eg `torch[gpu]` or `torch[nvidia]`.
We really need to move all AI/ML research off PyTorch to Candle or… just anything that isn’t Python or another old-gen, broken language like it.
That said I don't know what's wrong with using a Rust AI framework like Candle.
Weird turn of phrase very typical of AI.
The backbone is also a selective SSM rather than an LSTM controller, so the sequence half is closer to Mamba than to an NTM. DNC comparisons are fair too, and I should have cited both in the readme.
Edit: plausibly there’s no human in loop here at all, in which case fair play, I guess.
Why did you choose Rust? Why does that matter?
Nothing wrong with Rust. Lots wrong with the bandwagon that “in Rust” somehow adds value.
He could have said "... not written in Python" - would that have been acceptable?
There are a few layers of irony here, but see uv.
- https://docs.astral.sh/uv/
A non transformer language model written from scratch is more interesting than
“I did X in Python/Rust/Flavour-of-the-month-thing”