papersTODAY 04:00 UTC
TwinICL Benchmark Tests Multimodal In-Context Learning With Paired Counterfactuals
Researchers released TwinICL, a procedurally generated benchmark that provides matched text and image versions of the same tasks, allowing direct comparison of in-context learning across modalities. The paired counterfactual design is intended to isolate how much a model's few-shot performance depends on the input format rather than the task itself. The work appears on arXiv under cs.LG.