Design and operate the GPU cloud infrastructure that trains and serves Large Data Models at scale.
About NeoSpace
NeoSpace is an innovative startup shaping the future of technology with cutting-edge artificial intelligence solutions. We develop specialized AI models to optimize processes and transform our clients' experience. Our goal is to simplify people's lives and boost business efficiency by creating smarter, more accessible products and services.
What we're looking for
We're looking for a Cloud Specialist with strategic vision to work on the design, implementation, and maintenance of critical architectures in cloud environments, with experience in highly complex workloads.
If you believe that "infrastructure is code" and can move easily between different providers to support emerging technologies (such as AI), this role is for you.
Responsibilities
- Design and maintain robust solutions integrating AWS and OCI services, ensuring high availability.
- Develop, test, and version Terraform modules for 100% automated provisioning.
- Deploy and manage high-performance environments for AI processing, including NVIDIA clusters and low-latency interconnectivity.
- Implement IAM policies, secure connectivity (VPC/VCN Peering), and compliance controls.
- Monitor resource consumption and propose continuous improvements for performance optimization and cost reduction.
- Spread cloud culture and support development teams in adopting best practices.
Requirements
- Solid experience with core AWS services (EC2, S3, RDS, Lambda, EKS) and OCI (VCN, Compartments, Compute, Storage).
- Advanced knowledge of Terraform (preferred), AWS CDK, or CloudFormation.
- Deep knowledge of network protocols, VPNs, Direct Connect, and FastConnect.
- Hands-on experience with Docker and Kubernetes.
- Advanced Linux skills (troubleshooting and shell scripting) and experience with DevOps/SRE methodologies.
- Solid understanding of cloud security (KMS, IAM, encryption, and compliance practices).
- Sharp analytical thinking to handle complex incidents quickly.
- Ability to translate technical decisions for stakeholders across different areas.
- Genuine interest in testing new technologies and influencing the direction of internal products.
Nice to have
- HPC & NVIDIA: Experience with High Performance Computing clusters.
- Infrastructure CI/CD: Implementation of deployment pipelines (GitHub Actions, GitLab CI, or Jenkins).