图片/视频/文本/语音数据集,含标注
Annotated datasets for fine-tuning and evaluation across image-text, speech and video frames with split and labeling specs for ML teams.