AI inference is the process of a trained machine learning model running live data through its network to calculate an output, make a prediction, or solve a problem. While "training" is the phase where an AI learns and develops its capabilities from a massive dataset, "inference" is the phase where the AI puts that learned knowledge into practice in the real world.Here are the key aspects of AI inference:
- Real-time execution: The model applies its pre-learned rules to new, unseen data instantly.
- Pattern recognition: It identifies familiar structures in digital inputs like text, images, or audio.
- Decision making: The system generates specific outputs, such as translating words or identifying a face.
- Efficiency focus: This stage requires high optimization to run fast on consumer devices or servers.
Oracle Cloud Infrastructure's new X12 Standard Ax compute instances enable AI inference directly on Intel Xeon 6 CPUs, reducing cost per token by up to 6 times without needing GPUs. Learn more at the Oracle Blog.