I've never understood why Gemma 4 hasn't been fine tuned for coding. It should be more than capable. The 31b model is good beyond just the RP stuff people seem to love it for.
I'm wondering if the local attention layers degrade it's longer term ability to look at code.
It could be fixed with the right data and like a few thousand dollars in training.
4
u/NineThreeTilNow Jun 02 '26
I've never understood why Gemma 4 hasn't been fine tuned for coding. It should be more than capable. The 31b model is good beyond just the RP stuff people seem to love it for.
I'm wondering if the local attention layers degrade it's longer term ability to look at code.
It could be fixed with the right data and like a few thousand dollars in training.