Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Most of the work was getting CMake to find it. Just enable
RWKV_CLBLAST
and then drop the OpenCL & CLBlast distributions into the repository root like so:the actual folders after unzipping, of course!!
Marked as draft due to lack of testing—I unfortunately lost my bespoke chat script at some point and so can't really do my own experimentation immediately, but I do want to put this out there and have it available for others to see and test out for themselves.
Performance seems to be almost exactly on-par with CUDA in my experience. So maybe this will be getting CUDA-like performance out of Intel and AMD GPUs - exciting :D
It took me about 2 hours and 30 minutes of real time to complete this pull request :)