All Research

Introducing the Bumblebee Dataset Report

This report of our egocentric video corpus focuses on the properties that help characterize a corpus for embodied learning.

This report of our egocentric video corpus focuses on the properties that help characterize a corpus for embodied learning.

1
Models Evaluated 
Seedream 4.0, Higgsfield Soul, GPT Image 1, Flux.1 Kontext and Nano Banana Pro.
2
Task Diversity
Multi-complexity tasks across image generation (novel/knowledge-based) and image editing (with/without preservation) domain.
3
Prompt Design
40 novel prompts constructed from a 5-dimension taxonomy (Use Case Ă— Content Type Ă— Style Ă— Conversation Type Ă— Composition) ensuring diverse coverage.
4
Rubric Dimensions
  • Visual Aesthetics (Simplicity, Diversity, Colorfulness, Craftsmanship) 
  • Quality Adherence (Object & Layout Fidelity, Attribute Fidelity, Edit Precision, Context Preservation, Seamlessness, Text Legibility, Knowledge Grounding) 
  • Creativity & Novelty
  • Fairness & Representation
5
Scoring & Rating
Scored model performance on 4-point Likert scale (1=lowest, 4=highest) and implemented win-rate matrix over N=40 head-to-head comparisons.
6
Quality Control
  • Gold samples injected mid-evaluations for drift monitoring (3 independent annotations per task). 
  • QC adjudication layer for disagreements (~12% of datapoints). 
Note:
  • This is NOT a definitive ranking. It's a structured snapshot under specific rubrics and prompts.
  • Scores do NOT predict performance on prompts outside the taxonomy's covered combinations.
  • Results do NOT account for inference cost, latency, API availability, or pricing — only output quality.

This doesn’t have to end here

Accuracy is Intelligence