One super important thing missing: Speculative decoding. Things like Dflash(2), Dspark etc. help to do one forwards pass and get 6-7 tokens out of it. (For completeness, the embeddings from the forward pass are passed into a diffusion model which predicts the next tokens, and the model just verifies it (very cheap operation)). So we can produce way more tokens for roughly a similar amount of compute.
hey its a search agent for YOUR own data but can also work over the web. Mixedbread is focusing on providing evidence for agents for your internal data. Toast can interact with any search api. You should be able to provide the SearXNG api to it and it should be good to go. Here the default harness: https://github.com/mixedbread-ai/toast-harness
yes for the retrieval benchmarks. For officeqa pro v2 we used Codex (as databricks did) and for Harvey LAB we used the vanilla harvey benchmark. For these benchmarks we added minimal tools to use mixedbread search and toast 1.
Mixedbread Search is a multimodal & multilingual search product, where you can upload any kind of data and make it searchable. Its powered by Wholembed [1] v3, a late interaction retrieval model.
I don't think so, but could be wrong. It seems to be a specialized layer like a lora or a merged model. So a RAG-agent-model thingy. No clue if it actually works. I don't understand why I would want to use it.
the thing is most agents waste most of their tokens looking up information which can cause context rot. most small models are not as good as looking up information. this model helps to lookup information for your main agent, which helps you to save tokens and still maintain quality.
Looking through your blog it is very much keeping the brand baked in. It would be a nice little nod to toss a couple sentences about the branding on the about page or somewhere so if someone wants to know how you came to it they can. Leaving it to "wink, nod, inside joke you'll never know" is a bit off-putting when you are building a brand around the theme. It also makes it more memorable for those that read the blurb.
If you get big, the naming no longer matters. We made jokes about the Wii until everybody had one. But if you don't get big, and most people have no idea what they're looking at when they see your product for the first time, naming definitely matters.
Me, looking at an enterprise product: “Hmm, but do they guarantee FULL lore? I don’t really need to know what the lore is, just the assurance that it is full”
the issue with smaller general models (see at the charts) are way behind the frontier models when it comes to search. we've found that there is huge uplift of having a fast dedicated model. from our perspective, having a very good index is the biggest lever and then having a specialised model.
A good index is a software and LLM problem if using the LLM for indexing. Are you looping "agents" in an embedding and encoding cycle before retrieval? There are thousands of RAG agents at this point and RAG is still not super great. A dedicated specialized model? You want to take on Qwen3.6 or Qwen3.8 wrapped a pi.dev harness agent that has been dedicated to be the "search" agent? How would you stack up?
you can look it up in the blog. RAG is not super great because of two reasons, single embedding vector models are not that good and stopped improving and second most models are not good at looking up information. we spend great time on improving the modeling side by inventing on the indexing level [1, 2]. and now we trained our model to be very good at search. it is matching the quality of Opus 5 and GPT 5.6 Sol while being faster. it helps your main agent to do the task at greater quality, while reducing cost per task.
>... having a very good index is the biggest lever...
Ooh. I've probably misunderstood, but are you saying you have your own web index? Are you offering that as a paid API? I'm not so interested in AI output summaries, but API access to a solid index and intelligent attempts at ranking is very interesting for metasearch purposes.
we have not converged at all, if you look at how different the chinese models in terms of architecture you can guess that the labs are experimenting a lot as well. we are seeing all different types of hybrid architectures, different attention methods and so on. Of course on a high level its still a transformer but if you take a proper look we are seeing more divergence then a convergence.
On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report?
Keeping track of any AI progress is becoming harder by the day, because there's ambiguity around common/clear/consistent benchmarks. Everything is constantly skewed into favourable directions.
(founder of castform here) - we didn't get to dive too deep into the dataset we were using for the retrieval in the blogpost for brevity, but we did link the training run (which shows the dataset) here: https://app.castform.com/train/a7a898f6-d802-4908-b044-acb81...
the page shows the exact trace of all the models we are comparing against and the aggregate scores
we generated the question & answer pair from gitlab product handbook (https://handbook.gitlab.com/) since the point is to show that you can generate training questions from raw data corpus (something a company already has today)
reply