This report of our egocentric video corpus focuses on the properties that help characterize a corpus for embodied learning.
All Research
Introducing the Bumblebee Dataset Report
This report of our egocentric video corpus focuses on the properties that help characterize a corpus for embodied learning.

1
Models EvaluatedÂ
Seedream 4.0, Higgsfield Soul, GPT Image 1, Flux.1 Kontext and Nano Banana Pro.
2
Task Diversity
Multi-complexity tasks across image generation (novel/knowledge-based) and image editing (with/without preservation) domain.
3
Prompt Design
40 novel prompts constructed from a 5-dimension taxonomy (Use Case Ă— Content Type Ă— Style Ă— Conversation Type Ă— Composition) ensuring diverse coverage.
4
Rubric Dimensions
- Visual Aesthetics (Simplicity, Diversity, Colorfulness, Craftsmanship)Â
- Quality Adherence (Object & Layout Fidelity, Attribute Fidelity, Edit Precision, Context Preservation, Seamlessness, Text Legibility, Knowledge Grounding)Â
- Creativity & Novelty
- Fairness & Representation
5
Scoring & Rating
Scored model performance on 4-point Likert scale (1=lowest, 4=highest) and implemented win-rate matrix over N=40 head-to-head comparisons.
6
Quality Control
- Gold samples injected mid-evaluations for drift monitoring (3 independent annotations per task).Â
- QC adjudication layer for disagreements (~12% of datapoints).Â
Note:
- This is NOT a definitive ranking. It's a structured snapshot under specific rubrics and prompts.
- Scores do NOT predict performance on prompts outside the taxonomy's covered combinations.
- Results do NOT account for inference cost, latency, API availability, or pricing — only output quality.
.png)
%20(1).png)