Develop, deploy, and maintain LLM pipelines and RAG QA/search systems; design and optimize prompts and multi-agent LLM architectures; operate multi‑GPU/cluster inference; build evaluation pipelines for model quality, bias, and hallucination; collaborate with product and CS teams to integrate conversational AI.
Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.
We are seeking a highly skilled professional to join our team, focusing on advancing through innovative AI solutions.
The successful candidate will develop and refine Large Language Models (LLMs) to extract actionable insights, improve business decision-making, and optimize prompt design for more accurate outputs. Additionally, the role includes creating scalable and robust LLM/RAG frameworks tailored to customer service scheduling, fostering innovation and maintaining a competitive market edge.
This role is 100% Remote, Work from Home based.
Responsibilities
- Own the full LLM pipeline from data preparation to production real case usage.
- Design, iterate and optimize prompts (zero-/few-shot, chain-of-thought, tool-calling, etc.) to maximize model utility and safety across products and languages.
- Build and maintain Retrieval-Augmented Generation (RAG) QA/search systems that connect to multi-source knowledge bases.
- Familiar with vLLM/SGLang inference architectures and have proven experience deploying and operating LLM services on multi‑GPU or cluster environments.
- Design, implement and operate multi‑agent LLM architectures (e.g. LangGraph, CrewAI, AutoGen) including task decomposition, agent orchestration, memory sharing and tool‑calling workflows.
- Develop evaluation pipelines (automatic metrics & human feedback) to measure prompt and model quality, bias, and hallucination rates.
- Collaborate with product and CS teams to integrate AI models into conversational Chatbot in different scenarios.
- Track cutting-edge research, author tech blogs, and keep improve current architecture.
Requirements
- Master’s Degree or higher in Computer Science, Data Science or related field..
- At least 2 years of deep-learning/NLP experience, including 1+ year practical LLM work (SFT, DPO, RAG, quantization, inference optimization, etc.).
- Demonstrated prompt engineering & tuning expertise (few-shot design, structured prompting, prefix-/p-tuning, reward re-ranking, safety filtering).
- Practical experience building and deploying multi‑agent LLM workflows, with understanding of agent‑orchestrator patterns, shared memory, long‑horizon planning and guard‑rail design.
- Proficient in both English and Chinese communication for efficient cross team collaboration
Why Binance
• Shape the future with the world’s leading blockchain ecosystem
• Collaborate with world-class talent in a user-centric global organization with a flat structure
• Tackle unique, fast-paced projects with autonomy in an innovative environment
• Thrive in a results-driven workplace with opportunities for career growth and continuous learning
• Competitive salary and company benefits
• Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team)
Binance is committed to being an equal opportunity employer. We believe that having a diverse workforce is fundamental to our success.
By submitting a job application, you confirm that you have read and agree to our Candidate Privacy Notice.
Similar Jobs
Artificial Intelligence • HR Tech • Software • Generative AI
Remote annotator reviews Japanese text and images, draws bounding boxes, answers structured questions, and writes short summaries per guidelines to train AI systems. On-the-job training provided; assessment required.
Artificial Intelligence • Fintech • Software • Automation
Lead and mentor a Southeast Asia engineering team; ensure code quality and testing; drive project progress; define integrations strategy; design, build, and maintain third-party integrations; align engineering efforts with product and leadership.
Top Skills:
Aws ServerlessFastapiLambdaPostgresPythonReact
Artificial Intelligence • Marketing Tech • Sales • Software
Lead and scale the engineering organization for a regulated digital-asset custody platform. Own hiring, org design, engineering operations (on-call, incidents, release hygiene), delivery against roadmaps, audit and regulator readiness (SOC 2, SAMA, ISO 27001), and cross-functional alignment. Amplify an existing small, specialized team to meet regulatory and institutional requirements while preserving culture and retention.
Top Skills:
Aurora PostgresAWSBitcoinClickhouseEthereumGithub ActionsGoHsmKubernetesMpc/TssRustSolanaTemporalTerraformThreshold Signing ProtocolsTypescript
What you need to know about the Delhi Tech Scene
Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.



