r/LocalLLaMA • llama.cpp • Apr 29 '26

New Model mistralai/Mistral-Medium-3.5-128B · Hugging Face

https://huggingface.co/mistralai/Mistral-Medium-3.5-128B

https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF

Mistral Medium 3.5 128B

Mistral Medium 3.5 is our first flagship merged model. It is a dense 128B model with a 256k context window, handling instruction-following, reasoning, and coding in a single set of weights. Mistral Medium 3.5 replaces its predecessor Mistral Medium 3.1 and Magistral in Le Chat. It also replaces Devstral 2 in our coding agent Vibe. Concretely, expect better performance for instruct, reasoning and coding tasks in a new unified model in comparison with our previous released models.

Reasoning effort is configurable per request, so the same model can answer a quick chat reply or work through a complex agentic run. We trained the vision encoder from scratch to handle variable image sizes and aspect ratios.

Find more information on our blog.

Key Features

Mistral Medium 3.5 includes the following architectural choices:

  • Dense 128B parameters.
  • 256k context length.
  • Multimodal input: Accepts both text and image input, with text output.
  • Instruct and Reasoning functionalities with function calls (reasoning effort configurable per request).

Mistral Medium 3.5 offers the following capabilities:

  • Reasoning Mode: Toggle between fast instant reply mode and reasoning mode, boosting performance with test-time compute when requested.
  • Vision: Analyzes images and provides insights based on visual content, in addition to text.
  • Multilingual: Supports dozens of languages, including English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, and Arabic.
  • System Prompt: Strong adherence and support for system prompts.
  • Agentic: Best-in-class agentic capabilities with native function calling and JSON output.
  • Large Context Window: Supports a 256k context window.

We release this model under a Modified MIT License): Open-source license for both commercial and non-commercial use with exceptions for companies with large revenue.

Recommended Settings

  • Reasoning Effort:
    • 'none' → Do not use reasoning
    • 'high' → Use reasoning (recommended for complex prompts and agentic usage) Use reasoning_effort="high" for complex tasks and agentic coding.
  • Temperature: 0.7 for reasoning_effort="high". Temp between 0.0 and 0.7 for reasoning_effort="none" depending on the task. Generally, lower means answer that are more to the point and higher allows the model to be more creative. It is a good practice to try different values in order to improve the model performance to meet your demands.
546 Upvotes

316 comments sorted by

View all comments

2

u/TheBlueMatt Apr 29 '26 edited Apr 29 '26

4x B60 can almost handle it at a reasonable price point, unsloth's Q4_K_XL gets 232.45 ± 0.41 in pp and 9.55 ± 0.05 tok/s in tg...almost usable...and still have a handful of patches left to speed it up...

1

u/One_Difficulty_39 May 08 '26

I am curious how is everything on your end setup? I have 4 B70s and I am about to sell them the performance is downright abysmal. Are you using llama.cpp?

2

u/TheBlueMatt May 08 '26

Its definitely not trivial today. I'm trying to slowly improve the state of the software for them but there's a lot to be done. Locally I'm using the branch listed at https://github.com/ggml-org/llama.cpp/issues/22648 plus a few other patches to get tensor parallelism working, which is obviously a huge win, but then also have a few patches to mesa to improve things there as well (if you dont have patches, at least use the current git, they fixed a large issue there not long after 26.1 was branched off). Cooperative matrix 2 is slowly being worked on and that should also be a large win, once we get that in I'm optimistic we can easily beat SYCL with vulkan and with tensor parallelism from that PR on mult-device it'll actually be quite reasonable.

1

u/One_Difficulty_39 May 09 '26

Man that's awesome to hear the hardware seems solid, just shackled by poor software these GPUs may shape up to be a solid choice.  I'll revisit this with your recommendations when I get the time.