Choosing Pinecone's Next Brain
Comparing local models for Pinecone, and the GX10 settings that worked for us.
We were choosing the model that would do Pinecone’s next round of research. MiniMax was already our local baseline, and we wanted to see what Qwen and DeepSeek could do with the same kinds of work.
Pinecone is our ongoing research system. Its agents follow questions about things we’re interested in, from AI hardware and robotics to resources and emerging technologies. They gather evidence, keep notes, and return to earlier conclusions as new information comes in. A useful finding today should still be available to an investigation next week.
We run the models on a pair of GX10 computers. That gave us room to try large local models, but choosing one meant looking beyond the answer it wrote. We needed it to save the finding, preserve the source, update the right record, and tell us accurately when something remained unfinished.
Finding a useful configuration
Our first screen tried MiniMax, DeepSeek, two Qwen configurations, and GLM. Some ran out of test time, some had integration problems, and GLM never reached working inference. There was enough to identify promising candidates, but little reason to treat that first pass as a ranking of what each model could ultimately do.
We spent more time with Qwen Flash Next. We tried different reasoning settings, whether to keep its reasoning history between tool calls, and how to return feedback when an operation failed. A medium reasoning setting became our retained configuration. Turning reasoning up further didn’t consistently help, and structured feedback improved some checks while making others worse. We kept the existing text feedback.
That gave us a useful alternative to MiniMax. In one research summary, Qwen preserved a named source that MiniMax omitted. It also handled most of the tested actions correctly. Its source-based answers still added interpretations the supplied evidence didn’t support, so we needed to inspect the claims as well as the completed work.
These were small tasks in simulated workspaces. We could inspect the calls, read the resulting files, and compare what happened with what the model said happened. That last comparison exposed a more consequential MiniMax failure: it posted a message to the wrong simulated destination, then reported that it had used the requested one.
A tidy final answer could hide unfinished or misplaced work. That became something we looked for more closely as we tuned DeepSeek.
Getting the work finished
With reasoning turned off, DeepSeek left a recovery target unchanged while explaining the failure, then appended a status saying it was complete. Turning reasoning on improved that behavior. It repaired more targets and reported the remaining failure honestly, though it still missed the requested order of operations.
We settled on effort 25. Increasing it to 50 didn’t improve the overall tool or task-state results in that comparison and used more time and tokens. Then we tested additional instructions about following the requested order, recovering from the returned evidence, preserving unrelated content, and checking completion against the actual result.
That version finished the recovery tasks in the required order. I chose DeepSeek with those settings and instructions for Pinecone. Seeing it recover and complete the requested work gave us a useful basis for the choice, while source-based answers still needed scrutiny.
The settings we kept
These are the three configurations worth returning to from our comparison. The tool checks examine required operations; the task-state checks examine whether the requested work actually happened. We reviewed source-based answers separately, because changing a file correctly doesn’t tell us whether a research claim is supported.
| Model | Reasoning and history | Sampling¹ | Tool checks | Task state |
|---|---|---|---|---|
| MiniMax M2.7 AWQ 4-bit | Built-in reasoning; unchanged baseline | 1 / .95 / 40 / .05 | 33/36 | 12/14 |
| Qwen3.8 Flash Next NVFP4 | Medium; no preserved reasoning history; text feedback | 1 / .95 / 20 / 0 | 32/36 | 13/14 |
| DeepSeek V4.1 Flash EXL3 2.9 bpw | Effort 25; preserved reasoning history; tested instructions | 1 / .95 / 0 / 0 | 36/36 | 13/14 |
¹ Temperature / top-p / top-k / min-p. Qwen also uses presence penalty 0. Results are from related September 21–26 campaigns, not one simultaneous trial: Qwen's diagnostics were fresh held-out cases; DeepSeek's follow-up reused previously examined tasks. The action totals exclude two source-based workflows per configuration. These small sets don't establish long-running reliability.
MiniMax was the faster baseline. Qwen gave us useful synthesis and a working alternative. DeepSeek with the tested instructions gave us the strongest observed operation-order completion, which is why I chose it. All three still showed weaknesses in source-based work; the table doesn’t make any of them a flawless researcher.
The selected DeepSeek took roughly 1.8 times MiniMax’s summed task time on the common successful checks, observed on different dates with different cache conditions. That is one practical tradeoff to consider alongside the work it completed.
Copyable request settings
These are request fields for the pinned OpenAI-compatible servers. All three configurations require two GX10s with 128 GB unified memory each. The configuration download includes model revisions, launch settings, helper settings, and the complete DeepSeek instructions. Use the corresponding runtime; these fields aren’t interchangeable across arbitrary servers.
MiniMax M2.7 AWQ
{
"temperature": 1.0,
"top_p": 0.95,
"top_k": 40,
"min_p": 0.05
}
Qwen3.8 Flash Next NVFP4
{
"temperature": 1.0,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"presence_penalty": 0.0,
"chat_template_kwargs": {
"enable_thinking": true,
"reasoning_effort": "medium",
"preserve_thinking": false
}
}
DeepSeek V4.1 Flash EXL3 2.9 bpw
{
"temperature": 1.0,
"top_p": 0.95,
"top_k": 0,
"min_p": 0.0,
"chat_template_kwargs": {
"enable_thinking": true,
"reasoning_effort": 25
}
}
DeepSeek also needs the exact instruction overlay in the download and preserved reasoning between tool turns. Qwen uses text tool feedback without preserved reasoning history. These are the configurations we retained from our tests, rather than globally optimal defaults. Matching settings alone won’t reproduce the test harness or guarantee its scores.
We activated the rebuilt Pinecone system on September 27. Model selection gave us a starting point for the longer experiment: whether those agents finish useful investigations, keep the evidence behind their claims, and revisit conclusions when the facts change. That’s the work we’ll be watching next.