Learn Before
GPT-4 enables users to specify vision and language tasks in a manner parallel to text-only settings.
0
1
Tags
Prep Sessions
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.1 Foundation Model Capabilities and Benchmarking - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Visual Input and Multimodal Processing - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
When evaluating multimodal inputs across diverse visual domains, what type of output does GPT-4 produce?
How does GPT-4 allow users to combine text and images within a single input?
Transferability of Language Prompting Techniques to Multimodal Inputs
GPT-4 enables users to specify vision and language tasks in a manner parallel to text-only settings.
Discuss GPT-4's multimodal capabilities across diverse visual domains. In your response, identify at least two visual domains the model supports, state what type of output it produces, and describe how its performance on these multimodal tasks compares to its performance in purely text-based settings.