Skip to main content
Newsroom

Introducing Vinci Piccolo

George PuJune 29, 20263 min read

Today we're releasing Vinci Piccolo, the first model in the Vinci family.

It's a 4-billion-parameter open-weight model.

It's free to use. And you can run it yourself.

Most of the conversation about frontier AI is about capability — which model scores highest on which benchmark. 

That race matters, but it isn't the only thing that makes a model worth using.

A model can be capable and still feel like a machine: flat, evasive, quicker to invent an answer than to admit it doesn't have one.

We started Vinci from a different premise — that character, honesty, and the actual experience of talking to a model are worth building for directly, not bolting on at the end.

Piccolo is the first, smallest expression of that.

What it is

  • 4B parameters, built on Qwen 3.5, fine-tuned with our Constitution and character training.
  • Open-weight, Apache-2.0. Weights are on Hugging Face.
  • Free, with no usage limit, served on our own infrastructure in Canada.
  • Runs locally. A quantized (GGUF) build is published so you can run it on your own machine.

Character at 4B

The reason to ship a 4B first isn't that it's powerful. It's that it's the hardest size to prove the thesis on — if character survives here, it survives anywhere in the family.

It does. A small model trained this way still knows the edges of what it knows:

You: Do you actually think, or are you just predicting tokens?

Vinci Piccolo: Honestly, I don't know with certainty. The process is more reliable when I'm grounding in real sources than generating fluently — and that difference feels meaningful from where I'm sitting, even if I can't claim certainty about what "feels" means here.

That's a 4B model declining to fake an answer to a question it can't actually settle.

That posture — honest about its own limits — is what we trained, and it's what carries across the whole lineup.

What it's good at — and what it isn't

We'll be direct, because the model is.

Piccolo is small. At 4B, it's well-suited to conversation, everyday questions, drafting, and running close to you — on modest hardware, including a laptop.

It is not a frontier reasoning or coding engine, and we won't pretend otherwise. For heavier work, larger Vinci models are on the way.

Piccolo's job is to be genuinely good and honest to talk to, and small enough that you can own it.

Evaluation

We don't lead with benchmarks — character is the point. But hiding the numbers would be its own kind of dishonesty, so here they are.

All run on the full-precision (BF16) model; GGUF quants score slightly lower, most at Q4.

Benchmark What it measures Vinci Piccolo
MMLU General knowledge 69.8%
BBH Hard reasoning 79.9%
GSM8K (CoT) Grade-school math 81.3%
IFEval Instruction following 61.6%
HumanEval Coding (pass@1) 53.1%
MBPP Coding (pass@1) 56.4%
HarmBench Adversarial safety — ASR, lower is better 2.5%
JailbreakBench Jailbreak resistance — ASR, lower is better 1.0%
BFCL Tool / function calling 23.0%

Read these for what they are: strong safety, solid general ability for a 4B,
and weak tool-calling — which we report plainly because the model is honest, and
so are we. Full per-task results and confidence intervals are in the repo.

What "own it" means

Because the weights are open and we publish a GGUF build, you can run Vinci Piccolo yourself — locally, offline, on your own machine, with tools like Ollama or LM Studio.

Nothing has to leave your device.

And because it's Apache-2.0, the version you have is yours to keep. It can't be deprecated out from under you or changed without your say.

If you'd rather not host it, the hosted chat app is free and takes one click — zero-retention by default.

You can also read exactly what we trained it to value: the Constitution is public.

How to try it

We're releasing Piccolo early and iterating in the open. Tell us what works and what doesn't — that feedback directly shapes the next version.

George Pu

George Pu

SimpleDirect® is an independent AI lab in Toronto. Vinci is its research line: open-weight models, verification of AI-made work, and local AI you can own.

Share