Vision is useful when the answer depends on a diagram, chart, formula, or visual arrangement. It is also easier to incur a larger input charge than with a short OCR prompt. Start by deciding whether the image is actually necessary.
Image-capable choices in LuciQA
LuciQA’s Gemini options accept image input. Groq switches image questions to its Qwen vision model. OpenRouter’s two selected free Gemma endpoints accept images. Cloudflare switches image questions to Gemma when its text-only GLM model is selected. Cerebras likewise switches to its Qwen vision model. Mistral, OpenAI, and Claude are also wired for image input in the desktop application.
That is an application capability, not a promise that every account can use every model for free. Provider access changes. Gemini publishes model- and project-dependent limits; Groq has model-specific free-plan limits; OpenRouter’s free endpoints have shared account caps; Cloudflare measures its daily allowance in Neurons. Verify the exact model in your provider account before an important session.
OCR first or image first?
Use Ask from text (OCR) for clear paragraphs and multiple-choice options. Recognition runs on your device using English and Turkish OCR data. Review the recognized text if symbols, negation, or units matter. Use Ask from image / diagram when the visual structure carries the answer. The selected crop, not an entire automatic screen dump, is sent to configured providers.
In Pro Multi, review the suggested boxes. Each question is a separate provider request; selecting several questions and providers can consume quota quickly. A model’s confidence number is its own estimate, not a measured probability of being right.
See OCR versus vision and the free AI API table before choosing. This guide was last verified on 24 September 2026.