Job Description
We are looking for a Data Scientist to support the development of AI-driven features and applications across our product ecosystem. The role involves developing machine learning models, building APIs and AI services, contributing to Generative AI solutions, and assisting in model deployment and optimization. To succeed in this role, you should have a strong foundation in data science, hands-on Python programming skills, and a passion for solving real-world business problems using AI and Machine Learning while collaborating with senior AI leadership.
Responsibilities
- Contribute to the development of AI-powered features across multiple applications.
- Build, train, evaluate, and optimize machine learning models for business use cases.
- Assist in developing Generative AI applications using Large Language Models (LLMs), prompt engineering, and Retrieval-Augmented Generation (RAG) techniques.
- Design, develop, and maintain REST APIs and backend services to expose AI/ML capabilities using frameworks such as FastAPI or Flask.
- Perform exploratory data analysis (EDA), statistical analysis, feature engineering, and data preprocessing to support AI and analytics initiatives.
- Work with structured, semi-structured, and unstructured data to enable various AI use cases.
- Translate business requirements into practical, scalable AI/ML solutions.
- Implement data validation, feature engineering, and data quality checks to improve model performance.
- Support model experimentation, evaluation, hyperparameter tuning, and optimization using appropriate performance metrics.
- Assist in integrating AI models into production applications and support deployment activities.
- Monitor model performance, troubleshoot issues, and recommend improvements for accuracy, scalability, and reliability.
- Document experiments, model behavior, evaluation results, technical workflows, and deployment processes.
- Collaborate with AI/ML Leads, software engineers, data engineers, and product teams on solution design, model improvements, and system integration.
- Stay updated with advancements in Machine Learning, Deep Learning, Generative AI, and modern AI frameworks.
- Participate in code reviews, continuous learning, and internal knowledge-sharing initiatives.
Requirements
- 2+ years of experience as a Data Scientist, Machine Learning Engineer, or AI Developer.
- Strong programming skills in Python.
- Hands-on experience with Python libraries such as NumPy, Pandas, Scikit-learn, and data visualization libraries (Matplotlib, Plotly, or similar).
- Good SQL skills for querying, transforming, and analyzing structured data.
- Strong understanding of machine learning fundamentals, including supervised and unsupervised learning, feature engineering, model evaluation, cross-validation, and hyperparameter tuning.
- Good understanding of statistics, probability, hypothesis testing, and data analysis techniques.
- Experience designing and developing REST APIs using FastAPI or Flask.
- Familiarity with version control (Git), debugging, logging, and software engineering best practices.
- Exposure to Generative AI, Large Language Models (LLMs), prompt engineering, Retrieval-Augmented Generation (RAG), embeddings, or vector databases is an advantage.
- Ability to work independently on data science tasks while collaborating effectively within cross-functional teams.
- Good understanding of handling structured, semi-structured, and unstructured data.
- Strong analytical, problem-solving, and communication skills.
- Eagerness to learn and adapt quickly in a fast-evolving AI landscape.
Good to Have
- Experience with PyTorch, TensorFlow, or other deep learning frameworks.
- Experience with Docker and containerized deployments.
- Exposure to cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Familiarity with AI frameworks such as LangChain, LangGraph, LlamaIndex, or similar.
- Knowledge of vector databases such as FAISS, Chroma, Pinecone, or Milvus.
- Understanding of MLOps concepts, including MLflow, experiment tracking, model versioning, and CI/CD pipelines.
- Exposure to NLP, computer vision, recommender systems, or time-series forecasting.
- Basic knowledge of distributed data processing frameworks such as Apache Spark is a plus.
Education
Bachelor/Master degree in Computer Science, Data Science, Engineering, or a related field.