Why Pickaxe Should Host Open Weight Models

The Case for Open Weight Models on Pickaxe

Pickaxe gives us easy model switching, which is one of the best things about the platform. The list is missing an entire category, though. Every model on offer is closed weight, rented from a lab, at whatever price that lab decides to charge this quarter.

I want to make the case for adding open weight models, and eventually for hosting them on Pickaxe hardware.

Start with OpenRouter

The fastest path is an OpenRouter integration. One connection opens up Kimi, DeepSeek, Qwen, GLM, Llama, gpt-oss, and dozens of others, without Pickaxe negotiating a separate contract with each lab.

That alone would make Pickaxe the most model-flexible builder on the market. It also costs very little to ship.

The token math is not close

Open weight models are dramatically cheaper per token than frontier closed models. Frontier output pricing sits around $25 per million tokens. Strong open models land under $1, and some land closer to a dime.

That is not a 20 percent savings. That is an order of magnitude, sometimes two. For a platform that bills by usage, this changes what is possible to build.

OpenClaw is a token muncher

Agentic runs are where costs go strange. OpenClaw browses, executes code, and loops, so a single user session can chew through tokens that a one-shot form would never touch.

Running all of that on Opus is like hiring a surgeon to open your mail. Most agent steps are routing, parsing, formatting, and retrying. Kimi or DeepSeek or gpt-oss would handle those steps for pennies, with the expensive model reserved for the moments that actually need judgment.

Cheaper agent steps move OpenClaw from beta economics to production economics.

Privacy becomes a real answer

Right now, every prompt a Pickaxe builder sends leaves the building. That is a hard stop for anyone working with client data, medical information, legal documents, or unpublished manuscripts.

Self-hosted open weights let Pickaxe say something no competitor can say. Your data never leaves our infrastructure. That sentence closes enterprise deals.

Obscure models are a feature, not a footnote

The open ecosystem has models tuned for narrow jobs that the big labs will never bother with. There are models built for long-context retrieval, for structured extraction, for translation into specific languages, for classification at absurd speed, and for uncensored creative drafting.

Builders on this platform have use cases that a general-purpose frontier model handles poorly and a small specialist model handles beautifully. Today those builders have no way to reach them.

In-house is faster

Hosted inference on your own hardware means no shared rate limits, no noisy neighbors, and no waiting behind someone else’s traffic spike. It also means you can optimize for the workloads Pickaxe actually runs.

Speed is a user-visible feature. Every second shaved off a form submission is a user who finishes instead of bouncing.

Look at what OpenRouter actually sells

This is the part I would put in front of whoever runs the numbers.

OpenRouter does not train models. It does not run a research lab. It buys tokens, adds a small markup, and resells them. That thin slice over billions of tokens is the entire business, and it works.

Pickaxe already owns the hard parts of that business. There is a billing system, a credits ledger, usage metering, and a customer base that consumes tokens every day. The missing pieces are a price list and the metal underneath it.

Right now Pickaxe pays retail for tokens and passes the markup opportunity to somebody else. Every token a Pickaxe builder burns is margin flowing out the door.

Two revenue streams, not one

The first stream is internal. Match OpenRouter’s public prices, or undercut them, and Pickaxe keeps the markup on tokens its own builders already consume. Cheaper open weight inference means you can lower prices for users and raise gross margin at the same time. Those two numbers almost never move together, and this is one of the rare levers that moves both.

The second stream is bigger and it is brand new. Once the inference is running on Pickaxe metal, you can sell tokens to developers who will never build a Pickaxe. That is not a slice of the existing pie. That is a different pie.

At that point the markup stops being a markup. It becomes the whole price, minus electricity.

The ask

Ship the OpenRouter integration first, because it is quick and it proves demand. Watch which open models builders actually reach for. Then bring those specific models in house, where the margin, the speed, and the privacy story all live.

Short term, use OpenRouter. Long term, be OpenRouter.


Pricing check before you post: Claude Opus 4.8 is listed around $5 input and $25 output per million tokens. DeepSeek V4 Flash sits near $0.14/$0.28.