I strongly dislike CUDA. Once you have allowed that proprietary cr*p into your C++ codebase, it is very hard to get rid, and you end up with code that is either tied to a single vendor or an #ifdef hell, probably both.
The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually, like in Metal, OpenCL, and D3D12, etc.
These days we even have DSLs like Triton that make kernel writing much more ergonomic than anything you would hope to achieve in Rust.
> I strongly dislike CUDA. Once you have allowed that proprietary cr*p
Genuine question...why not just type "crap"? It's not even that much of a curse, but I've never really understood the point of self-censorship. If you don't want to curse then you could just use a non-curse word.
Because I know it is not technically crap, a lot of competent people worked on it, most with good intentions. I suppose it is better described as a cleverly designed Trojan horse than can infect your software and make that software become crap, in the sense that it becomes harder to maintain, increases code duplication, messes with your build system, ties your build system to platforms that have their toolchain binaries available, etc., etc., without bringing any long-term benefits over learning things the hard way.
The long term benefit is that there are more developers with CUDA experience available to hire than there are with any of the "hard ways" you mention.
Not saying you're wrong, but my career got a lot less frustrating when I started focusing more on the product and less on the ergonomics of the implementation. If you need to build a house and the customer isn't willing to pay for brick, you use vinyl siding.
For learning that may be a fine approach, but CUDA (in C++) really tries to hide what is going on behind the scenes, which is roughly:
1) code gets split between a host part that goes through your normal compiler, and a device part that goes through the GPU compiler. You may as well write the kernels separate and compile them via a separate compilation step, and keep your trusted host compiler for the host-side code.
2) data needs to move between the host and devices via explicit buffer transfers and synchronization steps, CUDA tries to hide this with annotated pointers, but it is really easier to think about those as just buffers that you allocate and transfer IMO, instead of trying to transparently share pointers between host and device like CUDA does.
3) kernel launches can we wrapped in a function similar to:
Instead of the funky <<< >>> syntax that CUDA for C/C++ imposes. The problem is that once you start putting that in your code, it stops being C++ and stops being portable to non-CUDA GPUs. The launching and grid settings can be a bit hard to grasp at first, but sugarcoating that in bastardized C++ syntax does not absolve from having to understand it eventually.
So a good place to start might be an OpenCL or Metal primer, depending on the hardware you have available. D3D12 (and probably Vulcan too) makes this much harder than it should be, with too much boilerplate but is overall a mature and well-designed API should you wish to develop for Windows. Starting with WebGPU might also be good these days. It has a very different shader language than the others, but the rest of the concepts are similar, and it has a strong emphasis on making things async, which is what you want for performance anyways.
Claude/Codex should be able to get you moving very quickly.
Thank you very much for the effort you put into your advice!! I think I will start with WebGPU (wgpu), even though I have an Apple Silicon Macbook. I would really prefer to work with Rust instead of C++ because I am not good with C++. (I believe) I am good with C, so my C++ code looks like C code, and I am kinda learning the differences as I learn CUDA, which is a terrible way to learn C++, I guess.
It doesn't read as emphasis to me. It reads like the person is trying hard not to curse, and they think "crap" is a curse word. It's a little bit adorable, like I'm reading a comment from an obedient child.
I guess you are not from the generation of texters. This how languages work we used to use * as a way to avoid getting censored it over time became a way to curse or give emphasis.
But typing "b*tch" with T9 is just as difficult (if not more) as typing "bitch". Anyway, I guess I never had a need to swear much over SMS. On IRC, on the other hand...
People are getting used to censor themselves in order not to be reported, banned, or «hurt » other sensibilities. The words « rape » couldn’t be written in instagram for example, what a great way to deal with such a serious issue. Mainly an American thing spreading away from young people if you ask me. Sorry America, just being honest here.
As much as I personally dislike TikTok, I don't think it is fair to it: cultural willingness for more sensor sheep on Internet started years before TikTok's popularity in the west.
It's not the sole cause, but I believe it's the main driver behind a bunch of specific substitutions that are mainstream now or nearly so. For example, dih, ahh, and unalive. They may not have been invented on tiktok, but that's where they incubated.
There's no such censorship on TikTok, it's entirely groupthink based on people saying "when I use that word my video is seen less so therefore it's being censored".
Youtube is a lot more guilty of it though, as well as demonetizing.
It sounds like there is a documented policy or proven that TikTok does it. Just people thinking it does leading them to self-censor. Then people see others doing it and copy it. So, I guess it is censorship but not by TikTok.
After reading through the threat here it seems more like a cultural thing.
The US has quite a lot of filters for profanity.
I remember from my youth that in 2009 Eminem was a guest in a Germany TV show and very happy to swear as much as possible without being censored.
https://www.youtube.com/shorts/2OC-yKZ5Yag
I think a string of non-alphanumeric characters would work much better here, like "Once you have allowed that proprietary @#$&% into your C++ codebase”
In any context I've seen, asterisks are for wrapping formatting and said formatting it to add emphasis. So being in the habit of typing `emphasised phrase`, for italics - regardless of whether the platform parses markdown/similar formatting, e.g. SMS.
To have an unclosed asterisk replacing characters in a word? I've only ever seen that as a way to bypass censorship. This spans communications from people currently in their 40s down to 20.
But this isn't perceived as emphasis at all. If I wanted to emphasize something, I'd be more likely to use something like *bitch* or something along those lines. Replacing a letter with an asterisk comes across as self-censorship, which is pretty silly - just use a different word if you're that uncomfortable with swearing.
Maybe more like p**p, as in "that cunt p**ped in my yard"?
I'll admit, it never once occurred to me that people might be using censored characters to provide more emphasis that a word is a swear, but I guess it does indeed do that, at least to the writer. Whether that comes across to the reader, and whether the writer cares that their intention was understood... I'm not so sure.
I was so disappointed when I tried reading tintin in other languages and found the dear captain was straight up using slurs in those. I wonder whether the english language ones have been edited over the years to remove that sort of thing
I certainly dislike how everyone on YouTube is saying “SA” and “unalive” and “corn”.
It’s one thing if it’s some funny commentary channel avoiding those words, but what bothers me is the true crime YouTubers. In the subject of true crime, rape and murder are just things that are probably going to come up, and when they refuse to use the appropriate language, it comes off as infantilizing, which is weird considering that my actual YouTube account is over 18, let alone the viewer using it.
I don't think those filters are even real, I think it's just mass-hysteria. I call these kinds of behaviours "traditions", but I'm not sure if there's a better term for it.
Basically someone comes up with something which is nonsensical, but plausible. Like believing that their videos are unpopular because they said the word "rape" and the algorithm magically got them, rather than because their videos suck. Then someone else sees that and starts thinking it is true. It silently spreads across the population.
I've seen this in organisations, where new recruits haven't been properly trained. Someone has come up with a method which is wildly incorrect and illegal, but plausible. The other new people around them have copied them. They've become slightly more experienced people, they've taught the next round of new people.
Before you know it, half of the organisation is doing something hilariously wrong, and they all sincerely believe it is the right way of doing it, because everyone does it. It's just self-reinforcing at that point.
> I call these kinds of behaviours "traditions", but I'm not sure if there's a better term for it.
In psychology that kind of thing is referred to as "superstition".
More specifically, "superstition" in this sense refers to the phenomenon of copying someone else's successful approach to a problem you have. (In your example, getting views on youtube.) Since you don't know what parts of their approach matter, you copy the effective parts and the ineffective parts equally.
I am sure they are bullshit. Like when they mute cursing and "risky" speech, but when you enable autogenerated subtitles they show up there. Youtube knows what thay said regardless if it's censored or not. It's so fucking stupid
A (baseless) hypothesis: perhaps there are plenty of YouTube creators who use the proper, mature terminology but you never see their videos because the algorithm really is penalizing them for it...
It's become so bad that even quality history youtube channels are frequently using euphemisms like "moustache-man" instead of just saying "Hitler", to avoid their videos being buried by The Algorithm, and therefore cut severely into their viewership.
> even quality history youtube channels are frequently using euphemisms like "moustache-man" instead of just saying "Hitler"
That can be quite confusing. You had German mustache-man, Russian mustache-man, French mustache-man (Petain), French small-mustache-man (de Gaulle), Spanish small-moustache-man (Franco)
It may be to bypass censorship, rather than self-censorship. Some platforms block or shadowban comments with curse words. Not sure about this platform.
I have written many words far worse than "crap" on this site. I haven't gotten in trouble over it yet.
I do find it a little amusing, because commenters stopped criticizing my cursing the moment I started getting a good chunk of karma here. I remember in 2016 someone criticized me for using the term "shitposting"...I don't think I've gotten that kind of criticism since 2016 though.
Your comment only makes sense in context if you believe in a deity who is too dumb to understand the difference between cr*p and crap. I for one do not worship a Bayesian spam filter.
> I've never really understood the point of self-censorship.
Some platforms disallow certain words. In order to bypass that, some people use the asterisks. That's just as one possible answer to your question; there can be many different reasons for self-censorship, but to me the most logical one is when one tries to work around crappy restrictions, such as on terrible reddit (they killed old.reddit recently; I retired before that due to moderators being insane, but I also said that if old.reddit is gone, I am gone anyway - the requirement to now log in, totally defeats old.reddit com's usecase. Then again reddit went downhill many years before that already, so not a real loss.)
IMO cr*p and crap are both valid but separate swear words. People have a wide option to choose from when they want to swear, and people like variety (much much more than LLMs do). People also tend to influence each other with their usages: cr*p is popular because it is popular.
Otherwise cr*p is just as good as crap, shit, horseshit, poopoo or such.
edit: * replaced with \* as HN interprets asterisks as formatting for emphasis. Thx latexr for informing me
"Daddy, what does cr*p mean?" Kids aren't stupid and this self-censorship isn't protecting anyone from anything.
(if a platform is serious about Bad Words for whatever reason (moral?) they would also forbid character replacements; ultimately it's the intent, not the word itself, that they try to steer with rules like that)
I guess I never understood censorship when it’s plainly obvious what you’re censoring. Anyone who can read will clearly know that it said “crap”, so I don’t see how it’s fundamentally different than just saying the word. You still put the word into my brain.
> Once you have allowed that proprietary cr*p into your C++ codebase
People have been doing that all the time for every kind of codebase. It's just part of the business. I don't see how it's worth having any emotions or opinions about it. Seems like you are wasting your energy.
Are win32 APIs proprietary? So you decide to use them, use a wrapper/UI framework, or don't develop for Windows. Easy choice.
Developing for embedded devices? So you read the manufacturers manual and implement based on the spec, use some sort of HAL if they are available, or you don't have a job. Even simpler.
Oh, does that mean I get to say you're ironic because, literally, they didn't tell anyone to do anything. They said they didn't understand the worth of the opinion. You're interpretation is selectively literal in order to be rhetorical.
Does that mean someone else gets say I'm being ironic because I'm selectively literal in order to be rhetorical? Well, okay, I guess it's harder now.
If you're going to make apps in windows, you need to call their proprietary API somehow. Maybe you do it via a wrapper library, or via electron or something. But that's the same thing, just with more indirection.
> If you're going to make apps in windows, you need to call their proprietary API somehow. Maybe you do it via a wrapper library, or via electron or something. But that's the same thing, just with more indirection.
Not even close to being true. You can invoke syscalls directly, just needs a bit of reverse engineering. I wrote a bare metal libc library, with (not a whole lot of) effort I'm fully able to interface with the kernel/open windows etc. Fully statically linked, no libc, no win32, compiled on Linux executed on Windows.
The problem is this isn't really well documented _at all_, and I even ended up attempting to get in touch with the Windows kernel dev team to give me the actual internal syscalls/endpoints, but they refuse to cooperate. Which is why writing anything for Windows is entirely pointless.
The problem is much deeper than that. Most OSes' syscall ABIs are not stable and could change without warning. What is stable is the dynamically-loaded libraries, shipped as part of the system. Linux is the notable exception here; the Linux kernel project doesn't ship a libc, and Linus is very famously opposed to "breaking userspace."
There's nothing that can stop you from using syscalls in theory, but if you want your app to be portable across different OS versions, past and future, you'd better not.
Incidentally, syscalls would also break Wine. The way Wine works is basically by shipping their own versions of Windows DLLs, which express their operations in terms of Linux APIs. Because Windows programs don't rely on syscalls, and call all system functions via the system-provided libraries, the Wine loader can just link Wine's version and let the program work normally.
> This isn’t even close to being true. Here’s a thing I did that made things way more complicated than is worth it for 99% of developers when there is a proprietary solution made so I do not need to worry about these things. Because it is so hard to work around it, it is entirely pointless to develop for one of the most used operating systems in the world.
Just being totally honest this is how I read this comment when I insert context that seems important to me. I respect having principles but at some point there needs to be more value in practicality over your codebase not being locked into a proprietary framework at all.
> Not even close to being true. You can invoke syscalls directly,
The windows syscall API is yet another proprietary windows API. Sure - you can call it without loading any DLLs. But you're still calling into a proprietary windows API.
If you really hate calling proprietary windows APIs that much, maybe stop developing for windows? Develop software for linux. Or make your own kernel, or whatever. But if you keep developing software for windows, stop fighting it. Unless you have a very good reason, your software should try to fit in on its host platform. It should behave well, and work like other windows software.
It's like travel. If you fly to France, try to fit in. Maybe learn a bit of French before you go. If you hate France, don't go.
The CPU on most machines is quite proprietary. I don’t understand this faux purity dogma.
Practical computing is not and never has been an abstract pure concept. It’s about making machines built by corporations to do usefull things at scale.
There is no ”non proprietary” computing unless you make your own stack.
Yes but there are business costs to using high-level proprietary tools and libraries. If you write your app using win32, you won’t be able to port is very easily. You’re also stuck with whatever bad or bizarre decisions Microsoft made.
It’s even worse for CUDA. GPUs are expensive, and now you’re vendor locked. You’re between a rock and a hard place. Either spend millions in engineering time, or millions on price-gauged hardware.
” If you write your app using win32, you won’t be able to port is very easily.”
This is wrong way around.
If you don’t support the platform your app runs on using the native api:s to the hilt your port is just bad.
If you actually want to support multiple platforms _you actually need to support_ them from the ground up.
This is speaking industrially and businesswise. A professional software business always has per-platform implementation resources. Or they have just one platform. Or they pretend they are multiplatform and then _everybody_ _daily_ fights with the problems this causes.
Obviously those elements that can be portable should be. It’s like Einsteins simplicity maxim - your codebase should be as portable as can be but not more.
” It’s even worse for CUDA…”
No these are just the business and market constraints. If this does not make sense for your offering then don’t use it. This feels like false FOMO - CUDA is not a silver bullet but it might be a specific solution to a specific problem.
Dead wrong. Win32 (externally) only seems stable, but internally it changes between Windows releases. Win7 syscalls are completely different from Win11 syscalls, meaning if I want to release a binary _without relying_ on Win32 I need to provide full syscall mappings _for each and every Windows version_. This doesn't happen on Linux.
That's literally the definition of it being stable. Programs written against an interface keep working despite the implementation changing. The Linux kernel also constantly changes internally but programs written against syscalls keep working, so it is stable; that fact doesn't stop being a fact just because I dislike perf_event_open(2) or whatever. This is all very basic and easy to understand.
I wonder which APIs you would use to port easily, because POSIX and Khronos aren't it either, as they are industry standards driven by companies where one has to pay for a seat at Open Group and Khronos offices.
Porting has never been hard. Just follow the platform guidelines. Make sane architecture. Done.
I mean _it's just work_. You don't need to invent anything. Just do the work.
What _is_ hard is when people run after silver bullets to avoid all this work.
Because people who don't understand software decide it would be cheaper to implement something only once. Or someone who does not really understand what they are doing insists that same C++ code runs automatically on all platforms.
AI has given the software engineers permit from the beancounters to do the sane thing.
Good software development orgs _have always_ done proper per platform ports.
Also - there is nothing wrong in supporting only one platform as such!
> Good software development orgs _have always_ done proper per platform ports.
I really wonder why this was never fundamentally fixed. How performant a certain instruction on a specific platform is, how well it is supported and potential equivalents or sets of other instructions to emulate an equivalent are usually all very well understood.
So there should be some graph of operations which can transform any software from and to the specifics of each platform. Especially because firmware + compliers + platform abstracting libraries are basically already just that graph, although (usually?) to lossy to be applied in reverse. Add the recent developments in very large scale statistics to it and it'd probably be quite possible to transform from and to generic intent in the implementation to the uniqueness of each platform. E.g. the theming differences between a MacOS UI and a terminal application served over serial or the processing capabilities of a VLIW CPU compared to a FPGA or a GPU server.
Considering the enormous amount of work that went into compilers, better debugging and intermediate representations it seems like a huge missed opportunity nobody seriously asked the question whether information could be emitted that would allow for decompiling all the way back to the generic intent.
That's a nice way of saying that it's a dependency clusterfuck.
I've never understood why we can't just expose the GPU ISA directly the way the CPU does. It's all getting compiled down at the end of the day so someone has to write a compiler for it either way. We'd be substantially better off IMO if it was all built directly into LLVM and then let middleware sort out the details.
That would require vendors to either stick with a single backwards compatible ISA like intel did for x86 or document how their graphics cards work.
CPUs manage this by changing the internal micro-architecture, but historically GPUs only needed to support a graphics API and used that abstraction layer to freely change the hardware.
> The CUDA runtime is a special case of one of the libraries provided by the CUDA Toolkit. The CUDA runtime provides both an API and some language extensions to handle common tasks such as allocating memory, copying data between GPUs and other GPUs or CPUs, and launching kernels. The API components of the CUDA runtime are referred to as the CUDA runtime API.
Launching kernels manually is an error prone PITA which I believe is the principle reason for CUDA's popularity. Having the compiler give an error when you mess up is a huge benefit. But having the compiler allow you to express "I want to launch this kernel over a grid with these dimensions, with these arguments" as a single expression is where the vast majority of the value comes from.
The having it all in a single file is mostly an artefact of the fact that it is C++, because C++ is single file at a time compilation. In D (which is multiple files in a single compiler invocation) with DCompute (which targets CUDA and OpenCL with upcoming support for Vulkan and Metal), you are required to write the kernels in a separate module, but you get all the benefits of the compiler complaining when you mess up _and_ the expressivity of "launch me this kernel".
Having also played with Metal and WebGPU (at least years ago), I would say that CUDA is, amazingly, the best GPGPU API we have. Do I wish we had an open source parallel programming language as good or better than it? Yes. But asymmetrically hating on CUDA like this is how we continue to lag behind it in UX.
> The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually
Not to mention that this is a completely sane way to use CUDA as well.
I know it's not the same thing because proprietary vs open software it's way less important but, generally if you are not ideologically against something you can easily follow the stream and do lot of nefarious actions, especially if the action has enough degrees of separations from the actual nefast outcome.
> The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually,
Yes, I also prefer doing it that way, but in Cuda with the driver API. Allows you to handle kernels like shaders, including editing and hot-reloading at runtime.
The reason I'm sticking with CUDA is because it's by far the most convenient API to use, without nonsense like 50-liners to alloc memory or the need to manage descriptors, bindings, queue families, etc.
> The reason I'm sticking with CUDA is because it's by far the most convenient API to use, without nonsense like 50-liners to alloc memory or the need to manage descriptors, bindings, queue families, etc.
I was there when the OpenCL committee was deciding on that sort of stuff.
As I recall, and it's been two decades and a lot of sleepless nights since then, there was real pushback at the time against OpenGL-style default bindings. So folks didn't want to establish an implicit command queue or any other default objects attached to other objects. Part of it is because OpenGL was perceived as clumsy and passé, some of it was because it is not friendly to multi-threaded applications.
Those first meetings were a shitshow full of tension, implicit threats from Apple, and backroom deals. Kudos to Neil Trevett for chairing the group; I I bet it wasn't fun for him either.
That's unfortunate. Cuda has shown that, when done right, defaults and a convenience layer can make for a well received API without sacrificing performance.
from what i can tell, you're going to be stuck with that no matter what you do
i'm currently using vulkan, and HLSL via dxc.
which should be portable but it's not.
apple refuses to support vulkan, and relies on moltenvk
and there's a bunch of OS/hardware/driver differences no matter what you do, that you'll probably have to feature test for, and compile a few different versions of your code no matter what you do
i think if you're doing something that you don't have to distribute to customers, just picking one stack and getting locked in has some appeal.
it leaves you vulnerable to lockin. but, especially in the age of ai, "claude, port this to vulkan" seems like a good enough defense against that
I don't mind CUDA, I do mind that all of the SDKs don't dynamically load the various CUDA shared libraries at runtime.. intertwining itself into your application linking process makes for extreme binary portability inconvenience.
> The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually
No. CUDA allows you to write all the code in a single file, and uses a preprocessor to split it back out and pass it through separate compilers, one for host and one for device.
This true, but you can write the two separately if you want.
The disadvantages of writing them together are listed in the various parent posts. But some code authors really like the convenience of having the two in the same file.
I highly recommend Julia for (scientific) GPU programming but it would be nice if there was a larger community and/or funding behind the GPU side of things. It has very few core devs for what it is.
Julia has had a great CUDA story for a few years now, and this about 9 days old. Rust rejects buffer aliasing at compile time using Rust's borrow checker, but shared memory in cuda-oxide currently requires unsafe, but then there's HuggingFace's Grout and mistral.rs, so yeah, Rust is picking up ground here on Julia. How is OpenCL's performance these days?
i'm surprised modular's doesn't get much traction. The promise seems super interesting, and chris latner has the record to back up his claims. If someone has an explanation..
Okay, so, GPUs are taking one more step towards being general purpose massively parallel machines. That's cool.
What would be even cooler though would be for GPU vendors to start giving us the user manual. An I mean the real user manual, that explains how to use their piece of metal when all you have is that piece of metal. That means a precise description of the wire protocols, the data format of the buffers we send to & get from the GPU, the ISA of the cores we have access to, the relevant performance characteristics…
In other words, enough information to write a state-of-the-art driver for any OS. That would be cool.
They don't release it because exposing a stable instruction set would kill their ability to quickly iterate, to release silicon with bugs that can be papered over with software fixes, as fixing bugs in chips is very expensive in terms of time to market, and undoubtedly to charge more for what looks like a hardware feature but actually is a software feature.
It's been this way for 25 years and I don't see it changing.
They don't need to do that to sell their hardware so why would they do that? On the other hand, they have strong incentives to not give you that level of access and information. The only way this will change is by having some disruption by way of a competitor who sells hardware with that feature as being a major reason why it takes off.
I think it would be much nicer, although unrealistic at the moment given the number of combinations of GPU vendors and variety of hardware, to directly target the underneath GPU ISA machine code.
Since I can write a simple compiler to target x64 machine code, it should be possible to write one to target my GPU.
Though, I'm certain that vendor lock is probably more profitable for them.
My thoughts on this, as a person who has recently coming around to working in this space is that up to now the convenience and "simplicity" of working in CUDA as it is has been a giant moat for NVIDIA. Having a whole toolchain with a C++ dialect and a giant extant pile of code out there that looked familiar to people meant they've "won" the AI wars.
And in that context NVIDIA had every motivation to keep their SDK somewhat abstracted higher up the chain and fully under their control and then be free to innovate in the lower bits. And this served them well as well as their customers.
My sense is that now with agent driven development this is basically evaporating. Agents are capable of at least prototyping/writing kernels for any hardware and ISA. e.g. OpenAI built their own custom hardware and ISA for it and then set agents loose on it writing kernels and claims great success. At least they're claiming this. And from my own experiences as a n00b entering this space, I can believe it.
TLDR I don't think vendor lock on the software side is going to work out for them as a strategy.
But luckily for them they continue to have really good hardware and good access to semiconductor fabrication. But just look at HotChips 2026 a couple weeks ago and look at the huge variety of new inference hardware coming down the pipe which looks completely unlike NVIDIA/CUDA.
Since NVIDIA owns huggingface now and huggingface has the excellent Candle [1] crate for inference on Rust, this seems like a good step towards nice native Rust kernels.
I will give you an outsider's perspective on an analogy in this case. It is easy to see Candle as a ML crate to use for neural networks in rust. I have used it, and it works well.
The analogy is Tensorflow 5-10 years ago. It is popular, and there are lots of material on it. You quickly learn from talking to people that due to whims, a collection of reasons, people's love of consensus that no one is recommending it; new people are not learning it. In this case, the Torch analogy is the Burn lib.
And to tie this back into GPU programming, Burn's backends use CubeCL, which lets you write compute kernels in a Rust DSL using #[cube], with a JIT compiler and autotuning machinery. It targets CUDA, AMD, Metal, Vulkan and WebGPU.
Nobody cares if kernels are written in Rust. Kernels were meant to be written in C, but if you want to go more high-level try Triton or a similar DSL that nicely abstract tile sizes etc.
kernels aren't meant to be written by any defined language. C is just a traditionally good default language that took over from assembly. No particular reason we have to stick with C.
SYCL is the natively polyglot counterpart, with practical implementations of it compiling down to the same sort of SPIR-V kernels as OpenCL. (OTOH, much of the current adoption on the open standards side seems to target the more widely supported SPIR-V compute shaders, via Vulkan compute.)
Not really, first of all it is for C++, not the range of languages supported by CUDA.
Before SPIR was a thing in OpenCL, Khronos could not understand why anyone would care about anything else other than C99, or why supporting Fortran on GPUs was at all relevant.
Secondly, from the competition only Intel cares about SYCL with their own sugar on top, OpenAPI.
AMD hasn't cared one second about it.
You may mention Codeplay, which is anyway an Intel owned company since 2022.
As for Vulkan, it doesn't have neither the features, nor the tooling that CUDA enjoys, it is the usual putting up with using LEGOs from different brands, with various pin sizes, that is so common with Khronos.
The momentum behind rust seems absolutely unstoppable at the moment, in the light of this, the adoption of Rust into the Linux kernel, and the adoption of for formally verified software by Amazon and Microsoft.
not too long ago, commenters on HN hated Rust and would never use anything written in Rust. So, now that Rust is in Linux, they shouldn't be using Linux either.
Really exciting but it reads like Claude instead of what Nvidia posts have generally been like in the past. I don't need nor want my tech blogs to sound like a young adult novel.
I’ve had this happen to me several time over the past weeks and it’s gone from quaint to humorous to farcical to outright “is-the-world-gaslighting-me” insane.
Just today I was reading Stanley Druckenmiller’s op ed in WSJ. This dude is like 80 and has made billions of dollars, and he got Claude to write his op ed???
Unbelievable. And the tells are so obvious, yet people still love the Claude-like quips and odd grammatical choices that read like halfway asshole halfway mid-sentence confusion.
That op ed was absurd. I respect Druckenmiller a lot and am always impressed with his lucidity in interviews. The Claude “ick” was all over his writing.
Yeah definitely Claude. Lazy authors, if you're going to get AI to write for you please use Astra instead - it makes way less annoying prose than Claude.
VectorWare founder here. We are working with them and stoked they are investing more in Rust. I just gave a talk at RustConf about our different takes (https://rustconf2026.sched.com/event/2KNQj/making-gpus-feel-...). The video isn't up yet but you should check it out when it is. The efforts are complementary.
Towards the end of the post, we (NVIDIA) mention that this work was done in collaboration with Vectorware and others in the Rust community. And we can't wait to build further with the community.
Anyone know when Rust's std::autodiff will become stable? Assuming this Rust support expands to other GPU vendors, autograd will probably be the only reason to use Slang instead of Rust anymore.
I was told in the 2025 LLVM dev meeting that it will always stay in nightly because it's not practical for them to provide long-term stability guarantees that is expected of stable Rust.
Not a cuda programmer, but since they’re making a new API, why would they already make it inconsistent at start? :-/ I’m referring to the examples a,b,c vs z,x,y (different ordering of output elements)
You write kernels in a Rust DSL using #[cube], it supports WebGPU through WGSL and Vulkan through SPIR-V, along with CUDA, AMD via ROCm, and Metal. (disclosure: I am a contributor)
I'm looking forward to trying these when they stabilize! I currently use WGPU for graphics, and cudarc for CUDA.
Note: Cuda-oxide is similar to Cudarc's host component, but uses a rust-style kernel dialect. Advantage: Share structs between host and device. Disadvantage: Trading standard Cuda kernels for a new, WIP dialect.
I haven't tried the tile API yet; looking forward to it.
The last time I checked, Cuda Oxide was Linux only, and required Async; these are why I haven't tried it yet.
cudarc been great for me, because it's easy to look up existing examples and references, and it maps 1-to-1 with what I see. I'm already having a tough time with CUDA itself, a dialect of it makes a tad harder to rely on previous work.
Seems more ergonomic in general though, both approaches they share, compared to cudarc, and less build infrastructure and fiddling with environments, which is great.
I think implication being organizations with 40,000+ employees and even more consultants and contractors plus a lot of budget are also using LLMs to draft public facing content instead of paying for content writers or even just proof readers .
It points to friction rather than cost economics. Same reason we are always surprised why multi billion dollar product companies with millions of install base prefer electron instead of a native app.
This does not imply that the organization is not paying for content writers or proof readers. It does suggest that they are not getting the value of paying for content writers or proof readers.
According to https://news.ycombinator.com/item?id=49417480, Anthropic hires writers who do not use LLMs, and reading their updates I also feel that they don't use LLMs for communication.
Pangram has an extremely low false positive rate. Even on adversarial examples.
One trade-off is even some obviously LLM text won't get detected by them, but they work really hard to ensure false positives are rare since a false accusation is much worse for society than someone getting away with LLM meatpuppetry.
I think you can't trust Pangram in a high stakes situation, but it is absolutely better than random noise at detecting AI-generated text. Which isn't surprising. If the distribution of probabilities can yield blatant Claudisms, it's not surprising it would also have more subtle deviations.
(Addendum: As I recall, LLM-generated outputs roughly follow Zipf's law, but the distribution still tends to have some subtle distinctions vs human text; pretty interesting, but I don't know where I heard this, so nothing to cite. Sorry.)
To be honest with you, I don't think I would be able to identify with high certainty that the bottom text is AI generated, so it definitely goes a long way to obscure the AI-generated nature of it, but I also think it still feels unnatural somehow. I realize my framing naturally calls into question whether I'm being honest, but I am being honest. Given my experience with similar "skills" (it's just chunks of prompt, nothing magical after all) I expected even less.
But still, this is all very strange because it wasn't that many generations of AI models ago that AI writing was a lot better - I'm talking GPT 4.1, Claude 4.5, that sort of era.
Anthropic newsroom posts on the other hand are carefully constructed and well-written in a way that I have not seen demonstrated by LLMs yet, past or present. I expect that they have well-paid staff who are careful with every detail of their public communications. When you put it that way, it almost feels unfathomable that they wouldn't, doesn't it?
I don't feel like either one of you really has a strong claim. "Doesn't appear to" is subjective, and of course it's impossible to prove one way or another.
You're simplifying the exchange a little too much. I said:
> Anthropic doesn't appear to use Claude for blog posts
My claim is literally the lack of evidence, which, yes, can't prove anything. This claim can be contested easily by showing evidence that they in fact, do appear to be using Claude to write prose in blog posts.
They said:
> Anthropic _absolutely_ does
Sounds pretty certain Anthropic is in fact, using Claude to write blog posts. Enough to emphasize "absolutely". That doesn't read like "I'm going off of vibes", that reads like "I can prove it". So, fine. Prove it. I don't believe it, and I want to hear the proof.
I'm skeptical, but it wouldn't be my first time being wrong. But flatly, if you make claims with this kind of certainty, yes I want to hear your proof.
My point in saying "Even Anthropic doesn't appear to be using Claude for blog posts" was not meant to be some striking revelation, I literally was considering it a prior to make another point. This on the other hand sure does sound like a striking revelation to me, that a lot of people across the Internet would be curious to hear. Like I'm sure these people would be interested:
I will admit that I am unnecessarily aggressive sometimes, but I wouldn't have changed my response much in any case. If you're going to make a strong claim like this, I want your evidence, not your vibes. Otherwise, the claim should be a lot weaker.
I also realize that this sort of brashness upsets HN a bit, but it is what it is. I pandered comments for votes in my 20s a bit, time to grow up, sometimes people won't like you. Sometimes I feel something deserves a brash response.
I don't really have any opinion on your tone; I just still don't agree with your framing. A lack of evidence would be neutral like "there's no evidence to indicate either possibility is more likely", but your phrasing conveyed that one possibility was more likely than the other. I pushed back against your follow-up because it seemed like you were arguing for a higher threshold of evidence than you provided.
Well, to be fair, you're correct. I am asking for a higher threshold of evidence. It's a stronger claim. I feel a stronger claim deserves stronger evidence.
McDonald's food is not even that unhealthy. I just tried a Burger King burger the other day and it's terrible. I think it's like 2000 calories in a single burger or something.
2k cal would be around 250ml of oil. or 350grams of peanuts. So doing with bread, meat, and other stuff alike requires over 600g of food, an excellent value to energy.
I get the impression that Nvidia employees don't care too much - I started seeing fully AI-written "documentation" on some of their smaller projects more than a year ago (i.e., before it was even slightly a good idea).
Yeah lol it’s basically the only reliable way to know how things work. Pre-AI, I read documentation for libraries that I used almost every day.
And now with AI I’m using it to fact check Claude. And still reading it for myself to understand why other peoples code is written a certain way. It’s basically the most important thing to reference when coding.
Sure today Claude can just read the library code and tell you what a function does or how to do something. But it still won’t tell you why something is a certain way or won’t figure out specifically-designed usage patterns as reliably as the author telling you “this is an example of doing x”
I really appreciated a friend reaching out to me with some PHP questions today. It was, to me, fairly basic but he was having a hard time grokking the documentation vs reading what his coworker wrote (some code using output buffering).
I brushed up on the docs since I haven't touched it in a couple years, explained my understanding of the ob_* functions, and gave him a very brief demo on a PHP playground.
He could have asked any LLM to tell him what that chunk of code did, and to explain the three functions, and instead he reached out to me. That felt _good_. Talking shop has always been a good way for me to form connections, because the pressure to socialize becomes task-oriented and you start to learn about how people think and feel, and that opens up easier paths for actual connection. It was nice.
Just like the Old Internet still exists - niche websites, mailing lists, probably a BBS or two (likely more right?), the pre-LLM world will trudge on, for a time. I hope LLMs actually lead to good things for people in the long run, and for now I personally will remain sparse in my usage of them.
The good news is your attention to actually reading and understanding documentation will differentiate you more and more as others (short-sighted, IMO) outsource understanding to an LLM.
when you start to internalize that these kinds of statements are an indication of how the average developer of the last 10-15 years operated the adoption rate of AI makes a lot more sense
Do you have the stomach to walk into a high school in the USA these days? Teachers use AI to generate assignments. Students feed the assignments to AI and submit the responses. Teachers feed the student submissions to an AI for grading.
As a high schooler going to a school with stricter rules on AI than most in my area, I can say that it's been going downhill ever since GPT 4. Teachers constantly use AI to create assignments(my French Teacher regularly handed us work with GPT 5.1 prose and emojis). Students are also rampantly using AI and bypassing school restrictions(We have a google account, making it easy to use Gemini if we just sign out), causing an inflation in GPAs and test scores. There's no easy solution to the problem, banning AI-tools only help somewhat as even typing into Google has AI web results, and students are quickly overcoming ways to restrict them. I have a friend that vibe coded an application that allowed his Mac Mini's desktop to be mirrored on his school chromebook, bypassing every restriction with sub 1-second latency. Of course, that opens the can of worms to whether schools should allow students to use AI...
> Of course, that opens the can of worms to whether schools should allow students to use AI...
We are starting to see results indicating cognitive decline due to AI in education, so no, we should do everything possible to ban it except for very limited fields.
LLMs aren't calculators or even computers, their generated output is too flexible, generic and basically starts replacing thinking.
Most likely they should only be allowed during late highschool years or just at university level, when people at least have a chance to learn how to research on their own.
I know many teachers who actually have respect for the profession, themselves, and the students. Thankfully that means they don't do this.
Whether this is a widespread macro trend is another issue, and would be terryfying.
If true, however, it would reflect on the values of the organization: we have spent decades underpaying teachers, and doing a poor job of pretecting schools from frivoluos lawsuits. Add into that, districts have thrown money into new buildings, have been suckered by Big Tech to adopt their policies (common core was pushed by Big Tech and has been a distaster as well as computers in classrooms). As a nation (the USA) we can't get our act together for a rigorous national exam, etc etc.
About half of the states in the USA require the ACT or SAT for high school graduation.
Alabama is one state that requires the ACT. The mean score in Alabama is below 18/36. Wisconsin is another. Its students score on average about 1 point higher than the national average of 19.4/36.
If you prefer states that require the SAT, the mean SAT score of students from Delaware is less than 980/1600, about 50 points below the national average.
I'll leave it to others to argue about whether these exams are rigorous.
I have two kids in engineering programs at a state University. They are allowed to use AI for homework assignments, but the homework is no longer worth any credit. They have a lot more papers, quizzes, and tests in class that count for their entire grade.
FYI, the world is a lot more decentralized than we think and even during the Dark Ages, guess what, that was happening in Europe and many places in the world were booming scientifically, technologically, etc.
Meh. It's not like we forgot how to make copper wires for landlines. We'll be fine. We'll live more or less like in 1880 or 1920, it's not a horrible life. I do hope we get to keep antibiotics, though.
It's true of most work in many and soon most white collar jobs, too. Claude writes some dense useless thing, everyone else uses Claude to summarize and write a reply to the thing. The Claude-submitted PRs get automatically reviewed and commented on by a GitHub Claude review bot. The programmer asks Claude to check out Claude's review comments to Claude. Claude pushes a commit to the branch and writes a comment. The Claude review bot reviews the commit and leaves a comment. The human [...].
My hot take is that it's not really that terrible in the long run for work since I think LLMs will probably be nearly or actually AGI and better white collar workers than most humans within 5 years of today. But it is very funny and surreal in the meantime.
It is definitely bad for school, though. Kids IMO should actually be encouraged to use LLMs but not in or for class work outside of an AI best practices class. Probably stop giving them homework (90% will always try to find a way to make AI do it) and have them solve problems in class hours with no electronics so that they're forced to not defer learning. This will become even more important once we have AGI.
Dude I am in slop fucking hell right now. There is still room for a human touch, without which the agents will lever us harder and faster into a world of incomprehensible garbage.
I totally concur. I'm almost lost for words at this stage. I need me some land to grow vegetables on and that's about it. Maybe some chickens. Every single day brings more despair (and not the prosperity we were promised).
It is the number 1 thing I cannot stand with Claude slop. It's a sort of anthropomorphization of language. Every "thing" does, produces, feels, wants, asks, answers, etc....
- "Launch is checked"
- "Question is asked"
- "The implementation answers"
- "The model wants"
- "The results name"
- "The connection surfaces"
- "The prompt wires"
- "The feature rides the mechanism"
Every single fucking thing is alive, wants things, and does things.
It's terrible. Infuriating. I want to rip my eyeballs out reading this filth. All. The. Time. "The anger is real".
Create any page with a file uploader. They all look the same now. It's like the Twitter Bootstrap days of responsive design. You'll get an icon which looks like ones on (on the drop space) those sites which are like "you must wait 60 seconds for this file to download".
It's so horrible. The human element has been completely removed and replaced by..... mediocre.
The human element has been completely removed and replaced by..... mediocre.
No it hasn't. The human element is still there, prompting the LLM. The change is that the human is happily accepting the first thing they get rather than critically looking at it and seeing a problem.
I don't think it's that because I see a lot of this in businesses where the human isn't paying the bill, or is even aware of what the bill is.
Humans are seeing either a shortcut to go faster (accepting low quality to move on immediately; reasonable if they're short on time) or a shortcut to lowering effort (accepting low quality because they don't care; not so reasonable but probably has a deeper root cause).
I hadn't read the article and read this comment as though NVIDIA themselves were implying that this library was checked but not trusted by them since it was fully LLM generated.
It gets fun when someone uses an uncensored model to bypass a refusal, but they accidentally pick one that was trained for erotic writing and brings its particular talent to the documentation task.
Yeah. I think if the text is written for other machines, then by all means have an LLM generate it, but if it is intended for a human audience, have a human being write it.
We are still much better at writing in a way that doesn't waste other people's time.
In this age of LLM written everything which has softly killed my motivation for learning Rust somewhat, this has revived my interest if not only for the fact the LLMs haven't yet been trained on this yet!
I've found sorta the opposite - in any area, it can just do everything for you, or it can be an incredible teacher. I've been re-learning a lot of higher-level math and it has been an knowledgeable, infinitely patient, always-available tutor. Of course, I could just have it do just about any math I want for me, but that's not the point.
Kinda the same with language/technology stuff - it can be a great tutor and it can scaffold other parts of a project for you. It can give you feedback and let you focus on the interesting parts.
I guess the motivation itself may be hard because of the fear of it taking over much of our jobs, but having this kind of help/feedback is pretty cool for the sake of learning things just because they are interesting!
> I've found sorta the opposite - in any area, it can just do everything for you, or it can be an incredible teacher.
Please don't. I've had all of Codex, Claude and Gemini convincingly tell me absolutely wrong stuff, pointing it out with easily verifiable example they come up with more and more weird reasons.
Things don't become correct simply because most sources are again - easily and logically verifiable - wrong. This already was a plague when people "just googled" stuff and effectively returned with the most SEO optimized answer. Now we have very convincingly written instances all over the place.
If these were singular instances I wouldn't be so worried, but if you are learning it already is very easy to learn something wrong. This is why back in the days when people still used physical books to learn new things it was a good idea to check first which books are actually recommended. There have been a lot of "experts" that wrote things they clearly misunderstood but worked for all the examples in their books.
To give a common example for both the backend and frontend devs, that isn't about a specific projects. LLMs and Google searches frequently turn out wrong results regarding CORS caching and how it works in relation to domains/hostnames. The circumstances under which Content-Disposition work are another example. I think a lot of wrong statements that LLMs are "convinced" about are due to wrong statements (sometimes in otherwise correct response) of popular Stack Overflow answers.
It's saddening how much wrong "common knowledge" exists in the industry. I have been bitten by a lot of these, but it feels when people don't even actually code and think anymore this will just rise forever.
Can these be wrong? Certainly. So can humans. Many of your examples are of humans being wrong. That doesn't make LLMs - or humans - useless. The fact that they are not infallible is not a reason to avoid using them and I'm not going to throw out a tool that has been incredibly valuable to me because someone on the internet got some bad CORS advice.
I was learning Rust slowly when the LLM enabled coding became good enough. I switched from learning to full on building with Rust. I still learn high level concepts as needed but I will not be able to write Rust on my own at all.
And that sounds scary but the way I got over the fear is by realizing there are many things that I do very well but I do not know their internals very well. Driving is an example. I barely understand what the steering wheel, clutch or brake pedals do. I have driven over 130,000 Kms and I will perhaps drive more than double that in the next many years.
I have been building software since PHP/Drupal days. Got into AWS S3 as a beta user. Adopted Memcached (and MQ) in 2008 out of necessity. Then Python/Django for 10 years. Then Rust. And tons of JS/TS. I owe a lot to my curiosity. I believe we can keep learning what we need and still delegate most of programming to agents.
There are two types of programmers: the pragmatists who see programming as a chore and would gladly never write a line of code again given the right tools, and the gardeners who don't want their enjoyable and rewarding garden-tending work taken away from them.
The "pragmatists" who get excited developing a prototype for a week before they realize they will never be able to ship something anyone else will use because each trivial change becomes exponentially more difficult for the LLM to implement and completely impossible for the "pragmatist" to reason about, with every new commit liable to break something else.
Still waiting for this revolution of amazing 10x software! It's been 10 months since Everything Changed in November, surely the 10x pragmatists could have leveraged their effective 8 years of development time? Or maybe we'll move the goalposts again and say that actually, Everything Changed with Astra, we'll just need to wait another three months?
I think OP’s point is that the payoff in learning a new language has diminished in this AI era. You can call that lazy, I’d consider it smart to consider whether you could be doing other, better, things with your time.
I had LLMs write a pile of cuda-rust code and they were quite competent at it. Ported a bunch of (C++) CUDA kernels over, and ground away on them til they got equivalent performance
what this article tells me is that no one at Nvidia actually cares about this project whatsoever. otherwise, they would have had a person actually write the announcement.
First of all, this is a pre-1.0 release that requires a nightly Rust compiler (if you choose the SIMT track with cuda-oxide) so that one is going to be unstable software.
Secondly, When an issue occurs with a kernel or you want to write your own custom kernel in Rust, now we need to diagnose if the problem came from either cuda-oxide (SIMT), Rust's side, CUDA or Tile (If you decide to choose the Tile track).
Another dependency into the list and course everything is open source except CUDA itself. So any issue that happens on the CUDA level, you are forced to wait for them to fix it.
Why? The point of a shader language is to define a function that outputs some graphics. Why should that not be portable between GPUs/CPUs of different vendors?
More generally any GPU computation, not necessarily graphics. The above argument could go that there is no point in high level languages for CPUs either and everyone should just always use assembly, which is obviously false. Same can go for GPUs.
Go's type system is not at the same level as C++, Fortran, Python, Julia, Haskell, Java, to quote the languages with CUDA support from NVIDIA and their partners.
The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually, like in Metal, OpenCL, and D3D12, etc. These days we even have DSLs like Triton that make kernel writing much more ergonomic than anything you would hope to achieve in Rust.
reply