> ML_LIBRARY // EXECUTORCH_v1.0
ExecuTorch
Meta / PyTorch Foundation — End-to-end solution for enabling on-device AI across mobile and edge devices for PyTorch.
client-edge-embeddedv0.3.0BSD-3-Clausequalified
Model Training
This library is a dedicated runtime engine for inference serving and does not train models.
Model Inference
Inference Accelerators:
CPUMPS
Deployment Targets:mobile, edge
Quantization:torchao INT4/INT8, XNNPACK FP16
What It Does
- +Execute PyTorch 2.x models natively on iOS, Android, and embedded DSPs
- +Delegate operations to Apple Core ML (MPS), Qualcomm Hexagon, and ARM Ethos NPU backends
- +Compact, modular C++ runtime with minimal memory footprint (<50KB for runtime core)
What It Does Not Do
- -Train models on-device (pure inference runtime)
- -Serve cloud multi-tenant concurrency
- -Run directly in web browsers
>Suitable Work Types
- Deploying mobile vision models (YOLO, segmentation) directly into iOS and Android apps
- On-device LLMs running on flagship smartphones (Llama 3 8B quantized)
- Wearable device AI with strict thermal and battery constraints
>Unsuitable Work Types
- Data center high-throughput model serving clusters
- Tabular financial forecasting
Data Residency Implications
100% on-device private execution.
Security Considerations
The .pte binary format is strictly verified at load time with zero dynamic code execution.
Operational Profile & Known Limitations
Maturity:emerging
Learning Curve:high
Ops Complexity:moderate
Cost Tier:free-oss
> Known Limitations:
- Exporting models requires strict compliance with torch.export dynamic shape constraints.
- Ecosystem is younger than TensorFlow Lite.
Associated Incident Patterns (Incidentpedia)
Enforce safeguards and monitoring to guard against these documented real-world failure modes:
> Primary Evidence & Benchmark Citations
ExecuTorch Documentationofficial-docs • >=0.2.0, <=0.3.x
2026-09-25HIGH
