A Plain-English Glossary of AI Image Editing Terms
AI image tools have accumulated a vocabulary that mixes machine learning research, photography and marketing, often in the same sentence. Some of the words describe something you can control, some describe machinery you will never touch, and some are mostly folklore left over from earlier tools. This glossary sorts them by what they actually affect, so you can tell which terms are worth learning and which you can safely ignore. Each group is explained in prose rather than as a wall of definitions, because these words make more sense in relation to each other than alone.
The words for what you type
A prompt is simply the text you write describing what you want. In an editing tool it describes a change to an image you have supplied, not a whole scene. A negative prompt, where a tool offers one, lists things you want kept out of the result. Negative prompts help but are less reliable than positive instructions, so they are best used for a few specific recurring problems rather than as a long list of exclusions.
A template or look is a prompt someone else has already written and tested, usually bundled with other settings, exposed to you as a single named option. The value is not that it is magic but that it has been tuned against many photos of a particular kind, so it behaves predictably. Foxi AI's named looks work this way, and they are built to preserve the real subject while the surrounding treatment changes.
A seed is a starting number that determines the random component of a run. Reusing a seed with an identical prompt gives you a similar result again; changing it gives a fresh variation. Not every tool exposes seeds. When one does, it is useful for producing near-variants of something you already like rather than starting over.
The words for the machinery
A model is the trained system that does the work. Different models have different strengths, training data and characteristic looks, which is why the same prompt can produce noticeably different results in two tools. Foxi AI runs on Replicate-hosted models from the FLUX family and gpt-image, which is the sort of detail worth knowing mainly because it explains why results have a consistent character.
Inference is the act of running the model on your input to produce an output. When someone says a run takes thirty seconds, they are talking about inference time. It is distinct from training, which is the far larger and more expensive process of building the model in the first place, and which happens long before you ever see the tool.
Diffusion is the technique most current image models use. In simple terms, the model learns to turn noise into an image by repeatedly removing a little of the noise, guided at every step by your prompt. Latent space is the compressed internal representation it works in rather than raw pixels, which is why the process is fast enough to be practical. Neither term changes anything you do day to day, but they explain why outputs vary between runs: noise is random, so the same instruction can take slightly different paths.
The words for size and shape
Resolution is the pixel dimensions of an image, such as 2048 by 1536. Megapixels is those two numbers multiplied and divided by a million, so the example is about three megapixels. Higher numbers mean more detail is possible, but only if the detail was actually captured. Enlarging a soft photo raises the resolution without raising the amount of real information in it.
Aspect ratio is the shape of the frame, expressed as a proportion: 1:1 square, 4:5 the tall shape common on social feeds, 16:9 wide, 3:2 the traditional photograph. Choosing the ratio early matters, because cropping a wide image to a square later means throwing away a third of it, and the part you throw away is often the part that made the composition work.
Upscaling increases pixel dimensions after the fact. Modern upscalers do more than stretch, they add plausible detail to make the enlargement look convincing. That is genuinely useful for a sharp image that is simply small, and it is not a rescue for a blurry one, because the added detail is invented rather than recovered. Foxi AI offers a separate upscaler for exactly the first case.
- 1:1 square, common for grids and thumbnails
- 4:5 tall, common for phone-first feeds
- 16:9 wide, common for headers and video frames
- 3:2, the traditional still photography proportion
The words for what goes wrong
An artifact is any visible flaw introduced by processing rather than present in the original scene. Blocky squares from heavy compression, halos along high-contrast edges, smeared texture where fine detail used to be. Artifacts in your source photo are a particular problem because the model may treat them as real content and amplify them.
Hallucination is the term for invented detail presented as if it were real: text that is almost words, a pattern that is almost the right pattern, an extra fitting on a product that never had one. It is not the model malfunctioning, it is the model doing exactly what it does, which is producing something plausible. It becomes a problem when you needed something accurate.
Drift describes the way a subject gradually stops being itself across successive edits, especially when each edit is applied to the previous output rather than the original. Banding is the visible stepping you sometimes see in smooth gradients like skies, where a continuous tone has been reduced to distinct strips. Both are easier to prevent than to fix: return to the original source for each new variation.
The words for control
A mask is a way of marking which part of an image should be affected, so the change applies inside the marked area and leaves the rest alone. Inpainting is editing within a masked region, typically to remove or replace something. Outpainting is the opposite, extending the image beyond its original edges by inventing what would plausibly have been there.
Strength, sometimes called denoise or edit strength, controls how far the model is allowed to move from the original. Low values change little and stay faithful; high values change a lot and drift more. When a tool exposes this, it is usually the fastest dial to reach for when results are either too timid or too far gone.
A reference image is a second picture supplied to guide style, colour or composition, distinct from the photo being edited. It is a way of showing rather than describing, which helps a great deal when the look you want is easier to point at than to put into words. Not every tool supports references, and where it does, the reference guides the treatment rather than replacing the subject.
The words for cost and speed
Credits are the unit most consumer AI tools bill in, rather than charging per second of compute. Different operations cost different amounts because they use different models and different amounts of processing. In Foxi AI, everyday looks cost about five credits and flagship ones about fifteen, with credit packs starting at 7.99 dollars for 250 credits and monthly plans available too.
Latency is the delay between submitting and receiving a result. Queue time is the portion of that spent waiting for capacity rather than actually computing, which is why the same job can feel fast at one hour of the day and slower at another. Most consumer image edits return in well under a minute.
A couple of adjacent terms worth knowing so you can recognise their absence: batch processing means submitting many images to be handled unattended, and an API is a programmatic interface for doing that from your own software. Foxi AI has neither, and it keeps one credit balance shared across web, iOS and Android, which suits interactive one-at-a-time work rather than automated pipelines.
Frequently asked
Do I need to understand diffusion to use these tools well?
No. It is useful context for why results vary between runs, but nothing about it changes how you should write a prompt. Time is better spent on being specific about what should change and what should stay.
What is the difference between resolution and quality?
Resolution is how many pixels there are. Quality is how much real detail those pixels contain. A large image made by enlarging a blurry one has high resolution and low quality, which is why upscaling helps some photos and not others.
Why do tools charge in credits instead of a flat fee?
Different models cost different amounts to run, so a heavier operation genuinely consumes more compute than a light one. Credits let that difference be passed through per action rather than averaged into one price.
Is a seed the same as a template?
No. A seed controls the random starting point of a single run, so reusing it produces similar variations. A template is a pre-written, pre-tested set of instructions that defines the look itself.
Try it on your own photo
Foxi runs this kind of edit in about a minute — upload a photo, pick a look or describe the change you want, and see the result before you pay for anything.