Skip to main content

> ML_LIBRARY // RAY_v1.0

Ray

Anyscale / UC Berkeley RISELab — Unified framework for scaling AI and Python workloads from training to serving.

distributed-computationv2.37.0Apache-2.0qualified

Model Training

Supported
Accelerators:
CPUCUDAROCMTPU
Distributed Training:Yes

Model Inference

Supported
Inference Accelerators:
CPUCUDAROCM
Deployment Targets:server

What It Does

  • +Distributed computing with lightweight actors and tasks
  • +End-to-end distributed AI orchestration (Ray Train, Ray Data, Ray Tune, Ray Serve)
  • +Zero-copy object store (Plasma) for memory sharing across workers

What It Does Not Do

  • -Implement low-level neural network operations natively
  • -Operate as a lightweight single-file script library
  • -Provide default authentication without enterprise wrapper configurations

>Suitable Work Types

  • Distributed LLM pre-training and fine-tuning
  • Large-scale distributed hyperparameter optimization with Ray Tune
  • Scalable multi-model inference serving with Ray Serve

>Unsuitable Work Types

  • Simple single-threaded local data scripts
  • Air-gapped edge microcontrollers
Data Residency Implications

Distributed cluster data movement; requires VPC isolation.

Security Considerations

CRITICAL: Isolate Ray Dashboard and ports. Unauthenticated access has enabled crypto-mining and data breach attacks (CVE ShadowRay).

Operational Profile & Known Limitations

Maturity:mature
Learning Curve:high
Ops Complexity:high
Cost Tier:high-compute
> Known Limitations:
  • Significant operational overhead in managing cluster topology and Ray jobs.
  • Object store memory pressure can cause actor eviction cascades.

Associated Incident Patterns (Incidentpedia)

Enforce safeguards and monitoring to guard against these documented real-world failure modes:

> Primary Evidence & Benchmark Citations

Ray 2.37 Documentationofficial-docs • >=2.20.0, <=2.37.x
2026-09-25HIGH