Skip to content

Resources · Answer quality

Enterprise search with a verifiable public evaluation.

Skyller scored 71.93 on EnterpriseRAG-Bench, a public evaluation with over 500,000 documents and 500 questions. In the leaderboard consulted on 6 September 2026, it ranked 4th among 22 systems, ahead of the evaluated configurations of OpenAI File Search, Amazon Q, Azure AI Search, Vertex AI Search and NVIDIA AI Blueprints. The results, method and answers are available for review.

See how it works
Skyller · Remote work policy
Skyller - AI AssistantRemote work policy

How many days can I work from home each week?

Skyller is an AI and can make mistakes. Remember to check.

Responding

Demonstration · illustrative data

71.93 POINTS

overall score of Skyller's EnterpriseRAG-Bench submission

+10.9 POINTS

difference from OpenAI File Search (61.03), in the configurations listed on 6 September 2026

20 OF 20 CASES

all cases in the no-answer category were correctly identified; this result applies to those cases

4TH OF 22

position checked on the public leaderboard on 6 September 2026; Skyller entered in 2nd on 12 August

The leaderboard

The configurations compared in the test.

The benchmark uses a shared corpus, 500 questions and evaluation method. Each row represents the submitted configuration of that system, including its models and settings. The results do not evaluate every product or version from each vendor.

Excerpt from the public EnterpriseRAG-Bench leaderboard on 6 September 2026, with each system's position and overall score.
metor.comPosition1Overall score80.34
Causal Dynamics LabPosition2Overall score78.95
TromlPosition3Overall score76.79
SkyllerPosition4Overall score71.93
OpenAI File SearchPosition8Overall score61.03
Amazon Q (Kendra)Position10Overall score48.96
Azure AI SearchPosition11Overall score48.42
Vertex AI SearchPosition14Overall score41.87
NVIDIA AI BlueprintsPosition16Overall score37.73

A selection of the top four entries and five enterprise search systems. The leaderboard listed 22 systems on 6 September 2026. New submissions may change rankings and results; see the full list on Onyx's site.

How to interpret the ranking

The ranking compares published submissions to this benchmark, which includes different types of systems. It is not a ranking of every product on the market or a certification of commercial availability. Skyller entered in 2nd place on 12 August 2026 and appeared in 4th place, with the same 71.93 score, in the 6 September snapshot.

The evaluation

An entrance exam for company assistants.

EnterpriseRAG-Bench is maintained by Onyx, which sells a product of the same kind and excludes itself from its own leaderboard to avoid a conflict of interest. The corpus, the questions, the grader and the leaderboard history are all public.

500K

A whole company, invented

The corpus simulates a technology company that does not exist: slightly over 500,000 documents across nine origins — internal messages, email, project tickets, shared files, CRM, meeting transcripts, code, support and wiki.

That is the volume of a real company, not a sample

ON PURPOSE

Messy like real life

After generating the documents, Onyx scrambles the corpus: files stored in the wrong place, near-duplicates across different systems, outdated versions and documents that openly contradict each other.

The test measures retrieval in disorder, which is everyone's case

500

Questions real people ask

Ten types, from the simple one with a single source to the one requiring an entire project, the one with conflicting information, and twenty whose answer simply is not in the corpus.

Including the question that breaks assistants: the one with no answer

INDEPENDENT

The grader is not Skyller

Each answer is compared with an expected answer and with the list of facts it must contain. Documents are classified by three independent evaluations, and the majority decides. The program and the rules are Onyx's.

No step of the grading goes through us

How the score works

Four measurements, not an opinion.

Correctness · 77.0
The grader reads Skyller's answer and the expected answer and judges whether they say the same thing. Style differences do not count against it; a figure that does not match does. Three out of four answers were judged correct.
Completeness · 79.14
Every question carries a list of facts the answer must contain, checked one by one. About four out of five expected facts showed up. The overall score of 71.93 combines this measurement with correctness.
Right documents · 81.6
Of the sources the evaluation deems necessary to answer, four out of five were found by retrieval. This is the measurement that says whether the right material reached the answer.
Noise · 8.86 per question
How many irrelevant documents retrieval brought along, on average. Lower is better here: every discardable file is noise the model has to filter out before answering.

The result

Type by type, no cherry-picking.

The overall average hides what matters most to a company: which question types the assistant is strong on. The test splits the 500 questions into ten types, and the table below is the whole of what Onyx published about Skyller.

Skyller's EnterpriseRAG-Bench score by question type, with the number of questions in each type.
Not in the corpusQuestions20Score100.00Right documents
Informal documentsQuestions20Score85.00Right documents100.0
Distant passages in one documentQuestions40Score83.33Right documents97.5
Direct questionQuestions175Score79.13Right documents88.6
Conflicting informationQuestions20Score74.57Right documents92.5
Question with a qualifierQuestions30Score72.56Right documents93.3
Big-picture questionQuestions10Score68.00Right documents
Indirect questionQuestions125Score64.47Right documents69.6
Documents from a whole projectQuestions40Score50.89Right documents64.0
Bring everything there isQuestions20Score32.08Right documents52.0

Source: Skyller's results file published by Onyx in the leaderboard repository. The dash appears where the evaluation defines no expected documents. The last two rows are the fronts the team is working on for the next submission.

What was tested

The submission used the product's public routes.

The 500 answers were generated through Skyller's public conversation endpoints, in the environment recorded for the submission. Its models and settings are part of that record; later product changes require a new evaluation.

  1. Starting point

    The literal benchmark question, sent in a brand-new conversation, with no category, no expected answer, no correct documents and no context from any previous question.

  2. What happened on each question

    1. 01The question came in through the product's own conversation route, with the platform's default model routing. The model that wrote the answer is OpenAI's — what the evaluation measures is Skyller's work around it.
    2. 02Retrieval ran across the 510,646 documents and returned the passages, through the same funnel that answers a customer conversation. On a 100-question check, the public search route and the assistant's internal path returned the same rate of correct documents (83.72%) and the same top document 93% of the time.
    3. 03Onyx did not settle for the answer file: it received access to the live product and sent 80 questions we had never seen, answered the same day through the same route, two days before publishing the result.
  3. What that supports

    The answer file kept here has the same fingerprint (SHA-256 dfe926a3…) as the file Onyx published in its own repository. Nobody has to take our word for it: the answers, the grader and the arithmetic are all open.

Skyller · Answer without a published source
Answer without a published source

What is the new rule that is still in draft?

Skyller is an AI and can make mistakes. Remember to check.

Responding

Demonstration · illustrative data

Where the test stops

  • The corpus is in English and belongs to a fictional company. The proof that counts for your company is a test with your documents.
  • The environment was tuned for this corpus. The result describes that configuration; quality with other documents, languages and settings needs to be measured.

What it means

Why this number matters for your company.

Identified all 20 no-answer cases
On the twenty questions whose answer did not exist in the corpus, Skyller correctly reported that it found nothing. This result does not guarantee error-free answers to other questions. The cited sources help the team check an answer before using it to support a decision.
It finds things in the mess
The test corpus was scrambled on purpose, with misfiled documents, near-duplicates and files that contradict each other. It is a portrait of the shared drive of any company with a few years of operating behind it.
The method and answers are public
Onyx maintains the benchmark, publishes the evaluation program and verified the submission with access to Skyller. The answers and results allow readers to examine what was measured and reproduce the evaluation.
You can verify it today
The leaderboard, the corpus, the questions, the grader and our 500 answers are all published. Anyone on your team can open them, compare with another vendor and redo the arithmetic.

When you decide

How to use this in your evaluation.

If you are comparing AI assistants for your company, three things this result makes easy to do.

  1. 01Whoever will use it

    Ask the same hard question of every tool you are testing

  2. 02Whoever decides

    Ask for the date and the source of any quality figure you are given

  3. 03Whoever runs IT

    Ask whether the test on offer ran on the live product or in a demo environment

What the evaluation showed

  • Finding the right documents in a large, messy corpus
  • Answering from them without contradicting the source
  • Saying it found nothing when the information isn't there

What only your own test shows

  • How it handles your company's documents and vocabulary
  • Who can see each document and each connection
  • Your team's routines running end to end

Questions

What people usually ask.

Get started

The proof that decides is with your documents.

A public leaderboard shows that retrieval works in a large, messy corpus, judged by someone who does not sell Skyller. What it does not show is your company. That takes one afternoon to find out, free.