Private beta

Source real agent trajectories, on your spec.

Open WebUI is where people actually work with AI: local models, hosted models, files, tools, agents, and private deployments. Open WebUI Data sources opt-in trajectories from that network, tool calls and outcomes attached, shaped to the request your model team writes.

We are onboarding design partners now. Early requests shape what gets sourced first, and what terms look like when this opens up.

There is no catalog and no fee to ask. You are charged only when we match your request to data you have reviewed and accepted.

Real work, not staged tasks
Coding, research, support, operations, analysis, and domain work, captured where it happened rather than performed for a fee.
Across every AI stack
Open weights, hosted models, private endpoints, custom providers, and local deployments all run through the same workspace.
Opt-in supply only
Open WebUI does not collect chats from installs. Contributors and approved organizations decide, request by request.
Buyer led
Start from the use case, chat shape, volume, labels, rights, and timeline your model team actually needs.
378M+
downloads
480K+
community members
150K+
GitHub stars

"Benchmarks show what a model can do. Real conversations show what it gets asked to do: the half-formed goal, the wrong file, the tool that fails, the correction, and whether the work ever finished. That gap is the data."

The sourcing problem

The data that would move your model is the hardest data to buy.

RL post-training, agent evaluation, and safety work need the same thing: real people finishing real tasks, with the context, the failures, and the outcome attached. An environment is only as good as the distribution it imitates.

Scraped text

Pages of prose with nobody completing a task. It cannot teach tool use, recovery, or intent, and it brings provenance questions your legal team has to answer later.

Annotation vendors

Six figures and three months for work written to a brief. Annotators perform the task they were assigned, which is why the failure modes never look like the ones your users hit.

Your own API logs

Only your surface and your models. The customers running the most interesting workloads are usually the ones on zero-retention terms, so you never see them.

Open WebUI Data is the fourth option. Real work from the workspace where it happened, sourced against your request, because none of it is collected by default.

The unit is the trajectory, not the message.

What a trajectory carries

The model and provider, the files attached, every tool call and what it returned, where the answer was wrong, how the person corrected it, and whether the task finished. That last field is what a grader has to score.

What that makes possible

Trajectories you can replay, outcome labels a reward model can be trained against, corrections that read as preference pairs, and held-out evals that were never on the public internet.

Environments and verifiers are yours to build. What is hard to come by is the usage distribution that makes one realistic, and the outcomes a rubric can be written against.

How a pilot runs

Start with a request, not a catalog.

Tell us the chats you want. We check whether opt-in supply exists. If there is a match, we scope a pilot around the data shape, review bar, usage rights, delivery format, and economics.

01
Write the request

Use case, domain, workflow, volume, languages, labels, exclusions, rights, and timeline. A useful request describes the model behavior you are trying to improve or measure, not just a row count.

02
We test for supply

We check the request against opt-in supply across individuals, communities, teams, and organizations. Some requests are too narrow, too sensitive, or not sourceable, and we say so early rather than late.

03
Sample before scale

You review data shaped like the real thing, against an acceptance bar you set, before anyone discusses volume.

04
Set terms and scale

You are not charged for the request, the search, or a match you turn down. Price and rights follow quality, review effort, scope, and timing, and they are agreed once there is data you want.

Request types and screening

Say what you want, and who it should come from.

A request has two halves: the workflow you need captured, and the profile of the people it has to come from. Screening is a request rather than an inventory, so we test every criterion against opt-in supply and tell you what is realistic before a pilot starts.

What to request
Coding and agent trajectories
Tool calls with their returns
Failure and recovery traces
Corrections as preference pairs
Domain expert reasoning
Safety, refusal, and red-team exchanges
Long-horizon research sessions
Low-resource and long-tail languages
Enterprise operations workflows
Held-out evaluation sets
Who to request it from
Country or region
Language the work is done in, including low-resource
Industry and sector
Role and function
Years of expertise
Licensed or credentialed professionals
Organization size
Deployment type
Number of distinct contributors
How the sample is balanced
Rights and provenance

Data your legal review can sign off on.

The hard part is not finding this data. It is showing where it came from, what you may do with it, and who was paid along the way.

Opt-in at the source

Open WebUI does not collect conversations from installs. Supply exists only where a contributor or an approved organization has turned sharing on, on terms they set.

Rights defined up front

What the data may be used for, how long it may be kept, and whether it may be passed on are written into the pilot rather than assumed after delivery.

Exclusions enforced

PII, secrets, regulated data, customer content, and any category you name are filtered or redacted before anything reaches your team.

Provenance you can defend

Where a trajectory came from, what was consented to, and what was removed. Documented for the auditability your GPAI obligations now require.

Contributors are paid

Every contributor sees the request, the reward, and the terms before sharing, and is paid for what they share. Workforce governance is part of your diligence, so it is part of ours.

Your request stays yours

Requests are not shared with other buyers, and scope is agreed per request rather than assumed. Open WebUI works for the request in front of it.

Buyer request

Tell us what you need.

For labs, model builders, data teams, safety teams, researchers, and enterprise buyers. Required fields take a couple of minutes. Screening and terms are optional, and only exist to make our answer faster. Prefer email? Write to [email protected] any time.

What happens next

We read the request against likely opt-in supply, then reply by email with a plausible pilot shape or a clear no. If it fits, we scope a sample first, and nothing is charged before you accept a match.

About you
Required
The data
Required
Contributor screening
Optional

Narrow the pool the conversations come from. Leave anything blank to keep it open, and we will tell you which criteria opt-in supply can actually meet.

How many different people the conversations come from.

A copy of your request is sent to your work email.

Open WebUI Data, answered.

Can we buy a dataset today?
Not off a shelf. Open WebUI Data starts from buyer requests: you describe the data you need, we test whether opt-in supply exists, and we scope a sample before anything larger.
Who is this for?
AI labs, model builders, evaluation and safety teams, academic researchers, and enterprise AI teams that need real human-AI conversations rather than staged tasks.
Why would Open WebUI have this data?
Open WebUI is the workspace layer people run on top of local models, hosted models, files, tools, agents, and private deployments. That is a view of real AI work that no single provider sees from its own API logs.
Do you collect chats from installs?
No. Open WebUI does not collect conversations by default. Supply exists only when a contributor or an approved organization chooses to share against a specific request.
When do we pay?
Only when there is a match you accept. Writing a request, having us test it against supply, and turning down something that does not fit are all free. Price then depends on quality, rights, review effort, scale, and urgency.
Can we request something highly specific?
Yes, and specific requests are easier to answer than broad ones. Tell us the workflow, domain, languages, labels, exclusions, and rights. If it cannot be sourced responsibly, we will tell you instead of stalling.
What happens after we submit?
We review the request, compare it against likely opt-in supply, and follow up by email with either a plausible pilot shape or a straight answer that it is not sourceable.
For contributors

Get paid for chats you already have.

If you use Open WebUI yourself, your existing conversations are the supply side of every request on this page. You see what is being asked for and what it pays before anything is shared.

Contributor page