Just want to say among many portable SIMD solutions I’ve seen recently (e.g. Fearless SIMD), this is the first that makes non-fixed vectors like SVE and RISC-V vector (RVV) easier to support. Glad to see they made this decision
How so? I imagine you'd still want to constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.
Mojo has an even more portable simd[1] type that isn't just generic over length but also over type. In my opinion it is almost always better to specialize for each platform and use portable implementation as fallback. It's a shame that just very few languages support Zig-like comptime, because it would be excellent for specializations without introducing runtime penalties.
> constrain the length to the maximum vector size supported by the lowest platform you want to support or you lose the portability and actually end up with code that performs much worse than the scalar alternative on some platforms.
Or, put a dynamic factor into your vector size and design everything around it. Such that every platforms can plug in their own factor and _scale_ the size of vectors. This is basically what LLVM IR does for SVE and RVV: `<vscale x 4 x i32>` where vscale is the said dynamic factor. Though the exact value of vscale is only known during runtime, it doesn't matter -- we still can design compiler optimizations and lowering around it. The generated binaries can then be portable across platforms with different vscale values.
I disagree, compiler optimization and software ecosystem in general already take a really really really long time to build, and I think it'll take longer without a standardized low-level binary interface i.e. ISA to enable rapid distribution.
And the approach you mentioned here:
> a family of ISAs that are ABI compatible such that one can compile down to a semi-pre-optimized portable IR, and just do the last bit per ISA
I think this is basically WebAssembly and PTX, and one may argue, Java bytecode. Yet look at how much efforts and time it took for WASM runtimes and JVMs to actually produce performant machine code (the "last bit per ISA" you mentioned) for just a couple of architectures! (e.g. X86 and ARM). And I wouldn't surprised if NVIDIA pour even more money on building optimization pipeline from PTX to each of their different uArchs.
one of the reasons I rarely read press releases is that I don't believe in promises -- I believe in _incentives_. In this case, what will Qualcomm be incentivized to do? What are in their interests?
Having Mojo support multiple platforms creates incentive to adopt Mojo and therefore write code in a language which can compile and run on Qualcomm hardware. This is good for Qualcomm.
However the danger is that the language sees wide adoption but nobody uses it with Qualcomm hardware. Instead it might encourage people to buy AMD. This is a terrible outcome for Qualcomm. They paid to boost someone else's sales.
So the incentive is to make sure it runs best on Qualcomm and to at least slightly hobble other hardware. But the safest thing overall is to support Nvidia, Qualcomm, and that's it.
I think a strategy for Qualcomm would be to use Mojo and Max as a software platform to drive AI inference on ARMv9 chips such as Snapdragon either at the edge (your smartphone) or in the cloud.
Ok, what will be Qualcomm incentive? Selling few hundred Mojo license for few thousand dollars each. Or making it open source hoping it may make big in AI / data science community and may help sell more Qualcomm hardware?
indeed, open sourcing is only half (or even less) of the picture: who is driving the open source community and how it is driven (i.e. governing structure) are probably more important IMHO. There are countless of cases where an open source project is either killed by slow death, or dictated by a single entity. Chris's previous projects like LLVM and MLIR are fortunate enough to grow and thrive organically, and that takes years if not decades to cultivate
reply