Senior AI / Software Engineer (Python, OCR, LLM & GPU Systems) (Remote)
The Senior AI / Software Engineer at Confidential Client will develop and maintain a large-scale document intelligence and AI platform using Python, OCR, LLMs, and GPU systems. This hybrid role involves building production-grade data pipelines, APIs, and AI model integrations across both on-premises NVIDIA GPU infrastructure and AWS cloud environments, focusing on scalable document processing, search, and AI inference capabilities.
About Careertakes
👉 Important disclosure: Careertakes is a third-party recruiting platform supporting this hiring process. If selected, you will be employed directly by our client, IT Services & IT Consulting; Data Infrastructure & Analytics; Software Development.
Applicants for this role may also receive access to additional matched opportunities through the Careertakes platform.
Overview
This role is for a Senior AI / Software Engineer focused on building large-scale document intelligence and AI systems. The position centers on Python-based production services, OCR and document processing pipelines, LLM integrations (cloud and local), and NVIDIA GPU inference infrastructure. This is a remote (telecommute) engagement with part-time hours and W-9 contractor terms.
Key Details
- Engagement: Part-time (approx. 20 hours/week). Opportunity to expand hours based on performance and project needs.
- Compensation: $20–$25 / hour.
- Work authorization: Candidates must be legally authorized to work in the United States and able to engage as independent contractors (W-9). Our client cannot provide OPT, employment/visa sponsorship, or other immigration sponsorship for this role.
- Location (authoritative): New York, NY (Remote / TELECOMMUTE)
Responsibilities
- Build production-grade Python services, REST APIs, and data-processing pipelines.
- Ingest and process large volumes of PDFs, HTML, scanned images, and structured/unstructured data.
- Design and implement OCR, extraction, parsing, normalization, and document-classification pipelines.
- Extract tables, entities, relationships, metadata, citations, and structured records from documents.
- Implement ETL and high-volume batch-processing systems for document workloads.
- Develop hybrid search solutions: embeddings, vector search, RAG, reranking, and evidence retrieval.
- Integrate cloud and hosted LLMs (e.g., OpenAI, Claude, Gemini, Qwen) as well as open-source models.
- Deploy and serve local LLMs on NVIDIA H100/H200-class GPUs with vLLM/PyTorch, CUDA, quantization, and batching optimizations.
- Build and maintain cloud deployments (AWS), including infrastructure for EC2, EKS/ECS, S3, RDS, OpenSearch, SQS, and IAM.
- Work with PostgreSQL, OpenSearch/Elasticsearch, pgvector, Redis, S3/MinIO, and queueing systems.
- Implement authentication, webhooks, Stripe integrations, and third‑party service integrations.
- Write automated tests for OCR/extraction pipelines, retrieval logic, and AI outputs.
- Debug complex backend, AI, GPU, and data-infrastructure issues and document architecture/implementation decisions.
Required Experience & Skills
- Strong Python engineering and backend fundamentals.
- Experience building production software and data engineering pipelines.
- Demonstrated OCR / document intelligence experience (processing PDFs, images, scans).
- Production experience with LLMs, embeddings, RAG, and vector search.
- Experience deploying open-source models to NVIDIA GPU infrastructure (experience with H100/H200 a plus).
- Familiarity with PyTorch, vLLM, CUDA, quantization, and GPU inference optimization.
- Strong AWS experience (EC2, S3, RDS, IAM, EKS/ECS, Bedrock familiarity helpful).
- SQL / PostgreSQL experience.
- Docker, Linux, and containerized deployment experience; familiarity with Kubernetes and CI/CD pipelines.
- REST API design and third‑party integrations.
- Ability to independently own and deliver on large, complex engineering problems.
Technical Stack (examples)
- Core: Python, FastAPI, SQL, PostgreSQL, Redis, Docker
- AI / LLM: Qwen, Claude, OpenAI, Gemini, Hugging Face, PyTorch, vLLM
- OCR / Documents: PaddleOCR, Tesseract, OpenCV, PyMuPDF, Docling (or similar)
- Search / Vector: OpenSearch/Elasticsearch, pgvector, vector DBs, embeddings, RAG
- GPU: NVIDIA H100/H200, CUDA, model quantization, local inference tooling
- Cloud & Infra: AWS (EC2, EKS/ECS, S3, RDS, SQS, IAM), Kubernetes, CI/CD
Ideal Candidate
- Deeply comfortable across OCR pipelines, Python data processing, and LLM integrations.
- Hands-on experience deploying and optimizing models on NVIDIA GPU servers (H100/H200).
- Strong track record of shipping backend systems that process thousands of pages and large datasets.
- Pragmatic about tradeoffs between cloud-hosted LLMs and on-prem / air-gapped local model deployments.
- Excited to work autonomously across data, infra, and model deployment boundaries.
Application Requirements (what to provide)
- Resume or LinkedIn profile
- GitHub and/or technical portfolio highlighting production systems
- Examples or short descriptions of production AI/data systems you built (OCR, LLM, ingestion, GPU deployments)
- Any specific NVIDIA GPU / model-deployment details (e.g., H100/H200, vLLM, quantization experience)
Qualified candidates will participate in a technical discussion and engineering review before engagement.
Equal Opportunity & Hiring Transparency
Careertakes and our client are Equal Opportunity Employers committed to building a diverse and inclusive workforce. We prohibit discrimination or harassment of any kind. To support a fair and efficient hiring process, AI tools may be used to assist with application review or resume screening. These tools do not replace human decision-making. Final hiring decisions are made by people.
If you have questions about how your data is used, please contact us directly.
Negotiate a higher salary! Check the salary ranges for this job type in your area.
View My Salary Range