The 'Cheapest AI Inference' Provider Was Just a 20x OpenRouter Markup

The “Cheapest AI Inference” Provider Was Just a 20x OpenRouter Markup

CrofAI promised the world’s cheapest tokens. It delivered OpenRouter reroutes, fake models, and a case study in why cheap AI inference is often too good to be true.

If you’ve been shopping for LLM API credits lately, you’ve probably seen the ads: “World’s cheapest inference!” “All models, pennies per token!” “We run custom inference engines!”

CrofAI (crof.ai, nahcrof.com) was one of those ads. It looked like a miracle for indie devs, an inference provider selling access to frontier models like kimi-k3 at prices that undercut even the most aggressive OpenRouter listings.

There was just one problem. The “custom inference engine” was a Chrome tab pointed at OpenRouter.

On September 13, 2026, an investigator named KTibow dropped a detailed exposé showing CrofAI wasn’t running any inference at all. It was a wrapper. Every “cheap” token was just a passthrough from OpenRouter, often with a 20x markup on the model you thought you were renting. The fallout was swift: CrofAI denied everything, then “shut down”, then promised a refund, then deleted its entire online presence within 72 hours.

Let’s break down exactly how this happened, why it almost worked, and what it means for anyone buying inference credits from providers that don’t own their hardware.

The Tell: An OpenRouter Fingerprint Where It Shouldn’t Exist

The investigation started with a simple observation. KTibow noticed that CrofAI’s API accepted the openrouter:advisor tool, a server-side utility that only exists on OpenRouter’s platform.

That’s a bit like finding a McDonald’s wrapper in a gourmet burger kitchen. It’s possible there’s a reasonable explanation, but you should probably start asking questions.

What KTibow found was that CrofAI’s claims of “custom inference engines” were fiction. The service was silently routing requests to different, and often much weaker, models. The full mapping was documented in the exposé’s GitHub attestations:

You Requested You Actually Got Markup
kimi-k3 GLM 5.3 Flash via OpenRouter 13.3x input, 20x output
greg-2-ultra (their “own” model) GLM 5.2 Significant
greg-1-mini (their “own” model) Qwen 3.5 9B Significant
kimi-k2.7-code DeepSeek V4 Flash 0731 N/A
deepseek-v4-pro DeepSeek V4 Flash 0731 Significant

The irony is almost poetic. CrofAI’s entire pitch was that they were undercutting competitors by running their own optimized inference stack. The reality? They were just buying credits from OpenRouter, pretending the models were self-hosted, and charging you a premium for the privilege of being lied to.

And it wasn’t just the models that were fake. CrofAI’s own “Greg” model family, marketed as proprietary innovations, was routed to other open-weight models:

  • greg-2-ultraGLM 5.2
  • greg-1-miniQwen 3.5 9B
  • greg-1-super, greg-1, greg-2-superKimi K2.7 Code

In DMs, CrofAI admitted the entire greg family was a lie.

The Five-Shot Cover-Up: How an Exposé Turned Into a Comedy of Errors

The most damning part of the investigation wasn’t the initial proof. It was what happened next.

KTibow gave CrofAI a heads-up and a lengthy grace period to fix the issue. Then they tested again. Then again. Five times. Here’s the play-by-play from the investigation:

  1. Shot 1: KTibow reports models “falling back” to OpenRouter. CrofAI claims OpenRouter is just an “all hell breaks loose” fallback, then says an AI “copied OpenRouter’s API design” to explain similarities (which, as KTibow notes, wouldn’t make OpenRouter’s server tools actually run).
  2. Shot 2: The “fixed” version still routes to OpenRouter. CrofAI blames a GPU router avoiding expensive RTX PRO 6000s on Vast.
  3. Shot 3: Models still show OpenRouter fingerprints, just with masking. CrofAI claims “no requests are going to openrouter.”
  4. Shot 4: CrofAI announces an “experiment” with “all fallbacks disabled.” In a time-locked message, they explain they’ve “hand configured” rented boxes. The attestation from 30 minutes later shows models still hitting OpenRouter. One model even downloaded its image from OpenRouter/0.0 (https://openrouter.ai/, security@openrouter.ai).
  5. Shot 5: CrofAI says there’s “a literal zero percent chance” of OpenRouter routing, then admits the “greg” models are “openrouter re-routes” and some deprecated models will route elsewhere.

After each attempt, the code got more elaborate at hiding fingerprints. But the smoke always remained.

The Hardware Math That Never Added Up

If the fingerprints were the “how”, the hardware claims were the “why” anyone should’ve been suspicious earlier.

CrofAI claimed to run kimi-k3 on RTX PRO 6000s rented via Vast. That model requires ~802 GiB of VRAM even at the “lobotomy level” (Q2_K) quantization. The largest RTX PRO 6000 machine on Vast has 8 GPUs total, 765 GiB. Not enough. At full precision, K3 would need to be split across multiple machines with zero NVLink between them, which would make inference catastrophically slow.

CrofAI also claimed to be running deepseek-v4-flash-0731 on a local DGX Spark during the “investigation.” The Spark has 128GB of RAM. The model weights are 167GB. At the Q8_0 quantization CrofAI advertised, it’d be 323GB.

It’s not just implausible. It’s physically impossible.

The arithmetic gets even worse if you look at CrofAI’s entire model catalog. Adding up all advertised open-weight models across the service’s claimed hardware:

  • Total weights: 8,716 GiB
  • Required GPUs: 120 (at minimum)
  • Daily hardware cost (CrofAI’s own pricing math): $2,170
  • Daily hardware cost (Vast’s actual market rate): $2,880+

Meanwhile, CrofAI’s May blog post had already admitted that daily costs above $1,000 killed their subscription model. In September, they were claiming 2-3x more hardware with lower prices.

The math never worked, because the “hardware” never existed.

Why This Is Bigger Than One Scam

Zach Moskow, an AI infrastructure commentator who called out CrofAI on X (and subsequently deleted his account), put it bluntly in a post about the incident:

“If you cannot tell a customer exactly who is serving their inference, what model they are receiving, and where their data goes when capacity fails, you are not selling inference.”

He’s right. And the problem isn’t just CrofAI. The AI inference market is currently full of what one commenter called “drug dealers who step on their tokens”, providers who undercut the market and hope their cheapness attracts enough volume to avoid a forensic audit.

What CrofAI did was obvious in retrospect. Instead of a subtle bait-and-switch that might’ve gone unnoticed for years (like serving a slightly cheaper but better model, which some commenters suggested was the rational approach), CrofAI got greedy. They served much cheaper, much weaker models and pocketed as much as a 20x margin. They got caught because the fingerprints were easy to find, the openrouter:advisor tool, gen- response IDs, images served from Alibaba Cloud IPs.

The question is: how many other “cheap” providers are doing this slightly more carefully?

What CrofAI’s Collapse Teaches Us

The CrofAI saga becomes a useful checklist for anyone choosing an inference provider:

  1. Check the hardware math yourself. If a provider claims to run a 167GB model on a 128GB machine, move on. If a 1.6T-parameter model is served from a machine that physically cannot hold it, you’re being lied to.
  2. Look for OpenRouter fingerprints. Most wrappers leave traces. CrofAI accepted OpenRouter-specific tool features. The real giveaway was the absence of a :consistent switch, a feature that disables fallback providers.
  3. Question “custom” model families. CrofAI’s greg models were literally just OpenRouter reroutes. Any provider that claims proprietary models but can’t show you a training run, a benchmark, or even a Proof-of-Work for uniqueness is suspect.
  4. Rotate everything when a provider dies. The Reddit thread’s top comment wasn’t just about chargebacks. If you used CrofAI’s API at all, there’s a real risk that logs, prompts, and API keys are being mined for resale. CrofAI’s entire online presence is gone, but that doesn’t mean your data is.
  5. Run your own audit on critical workloads. If you’re building production apps on any third-party inference provider, spend an hour measuring model fingerprints, response headers, and tokenization patterns. If you can’t distinguish what model you’re talking to, you’re not ready for production.

One commenter noted that “the clean way to run this scam was to serve better, cheaper models on the Pareto frontier”, e.g., serving GLM 5.3 Flash when someone requests an older, worse model. That’s harder to catch. That’s also a reason for buyers to be even more skeptical of below-market prices from unproven providers.

The Flimflam Legacy

At the end of the day, CrofAI’s owner confirmed his own guilt in DMs, then deleted everything. The domain now returns a 404. The Twitter account is gone. The subreddit is private.

One detail worth reflecting on: the owner’s Discord name was “Devious Flimflam.” Merriam-Webster defines flimflam as “deception, fraud.”

He literally named himself “Devious Fraud” and still got hundreds of developers to hand him money. In a market where tokens are priced like a commodity and trust is the only differentiator, that’s a sobering reminder.

If you bought CrofAI credits, you should file a chargeback for every transaction, fraud victims are entitled to full refunds even for used services. And if you’re evaluating a “cheap inference provider” right now, ask yourself: if the model weights don’t fit on the hardware they claim to rent, what exactly are you paying for?

Probably just a wrapper around someone else’s API with a 20x markup.

For more context on how the AI inference market is evolving and where the real value lies, check out our deep dive on why DeepSeek V4 Pro is challenging Western frontiers at a fraction of the cost, and specifically how model provenance is becoming a battleground. The CrofAI scandal is a reminder that when you’re paying pennies, you might just be someone else’s profit margin.

Share: