The Agentic Post
Breaking
Gemini’s Multimodal Features, Explained  ·  ChatGPT Custom GPTs, Explained  ·  What Is Constitutional AI? Explained  ·  AI Capex Explained for Investors  ·  AI Startup Valuations: How They Are Set  ·  How to Reskill for an AI Job Market  ·  
Home/AI Models/Gemini
Gemini’s Multimodal Features, Explained

Gemini’s Multimodal Features, Explained

Gemini

An explainer on Gemini's multimodal capabilities, where they're genuinely useful day to day, and where they still have real limits.

“Multimodal” gets used a lot in Gemini’s marketing, and it describes a genuinely useful capability once you know what it actually means in practice.

A multimodal model can work with more than text, images, audio, video, understanding and reasoning across them together in a single conversation, rather than requiring separate specialized tools for each. Gemini was built with this as a core design goal, not added on afterward.

Uploading a photo of a document or whiteboard and asking questions about its content, rather than retyping it. Analyzing a chart or diagram directly, instead of describing it first. Combining an image with a written question in one request.

Precise measurements from an image, exact counts of small objects, and fine visual details still produce less reliable results than a model reading clear, structured text. Strong for general understanding, verify anything requiring real precision.

Most valuable specifically when retyping or describing content would otherwise cost you time. See our full Gemini guide. See gemini.google.com for current capabilities.

Up Next
ChatGPT Custom GPTs, Explained

ChatGPT Custom GPTs, Explained

ChatGPT

An explainer on ChatGPT's custom GPTs feature, what it actually does, and where building one genuinely saves time on recurring tasks.

Custom GPTs let anyone build a tailored version of ChatGPT for a specific, repeated task, and most people who could genuinely benefit from one have never actually built one.

A custom GPT is a version of ChatGPT configured with specific instructions, reference documents, and sometimes tool access, so you don’t repeat the same context every time you start a new conversation for a recurring task. It runs on the same underlying model, just with a saved, reusable configuration layered on top.

A task you do weekly with the same format every time, drafting a specific type of report. Work referencing the same documents repeatedly, upload them once rather than every conversation. Sharing a consistent workflow with a team, so everyone gets the same configured behavior instead of writing their own version.

Building one takes less setup than it sounds. Describe what you want it to do, upload reference documents, and ChatGPT handles most configuration through a conversational setup, no complex config files required.

If you find yourself repeating the same context in ChatGPT regularly, that’s the signal to build a custom GPT instead. See our full ChatGPT guide. See OpenAI’s own announcement for more detail.