> ML_LIBRARY // RAY_v1.0
Ray
Anyscale / UC Berkeley RISELab — Unified framework for scaling AI and Python workloads from training to serving.
distributed-computationv2.37.0Apache-2.0qualified
Model Training
Accelerators:
CPUCUDAROCMTPU
Distributed Training:Yes
Model Inference
Inference Accelerators:
CPUCUDAROCM
Deployment Targets:server
What It Does
- +Distributed computing with lightweight actors and tasks
- +End-to-end distributed AI orchestration (Ray Train, Ray Data, Ray Tune, Ray Serve)
- +Zero-copy object store (Plasma) for memory sharing across workers
What It Does Not Do
- -Implement low-level neural network operations natively
- -Operate as a lightweight single-file script library
- -Provide default authentication without enterprise wrapper configurations
>Suitable Work Types
- Distributed LLM pre-training and fine-tuning
- Large-scale distributed hyperparameter optimization with Ray Tune
- Scalable multi-model inference serving with Ray Serve
>Unsuitable Work Types
- Simple single-threaded local data scripts
- Air-gapped edge microcontrollers
Data Residency Implications
Distributed cluster data movement; requires VPC isolation.
Security Considerations
CRITICAL: Isolate Ray Dashboard and ports. Unauthenticated access has enabled crypto-mining and data breach attacks (CVE ShadowRay).
Operational Profile & Known Limitations
Maturity:mature
Learning Curve:high
Ops Complexity:high
Cost Tier:high-compute
> Known Limitations:
- Significant operational overhead in managing cluster topology and Ray jobs.
- Object store memory pressure can cause actor eviction cascades.
Associated Incident Patterns (Incidentpedia)
Enforce safeguards and monitoring to guard against these documented real-world failure modes:
> Primary Evidence & Benchmark Citations
Ray 2.37 Documentationofficial-docs • >=2.20.0, <=2.37.x
2026-09-25HIGH
