Data Scientist
Job Description:
About Outsentia
Outsentia is a global talent and operations partner helping companies build high-performing offshore teams in regions like Pakistan. From sales and customer success to engineering and research support, we help our clients find top talent and scale quickly, without compromising quality. Our mission is to bridge exceptional global talent with fast-growing companies that need flexible, reliable, and cost-effective support.
We are currently hiring on behalf of one of our clients, a U.S.-based research platform that leverages expert insights to produce institutional-grade research for financial services firms around the world.
Data Scientist (Python, Pandas)
Location: Islamabad, Pakistan (6 PM to 3 AM) (3 days onsite)
Reports To: Head of Product
Client: U.S.-based Research Platform
Role Overview
We're seeking a Data Scientist to drive analysis, insights, and data workflows on a modern research platform hosted on AWS. The company has a wealth of data — including transcripts from expert calls and public-facing data — that we want to dive into to surface trends, so the ideal candidate has strong experience with Python — particularly data analysis libraries like Pandas — and can design and develop end-to-end data processing pipelines for both structured and unstructured data, including extracting features from text and normalizing semi-structured data into a centralized data lake. Once that data is organized and documented, you'll partner with AI agents built in-house to run freeform, exploratory analysis and uncover trends across the platform's data. You'll work closely with full-stack developers and product stakeholders, and you'll have significant latitude to shape analytical methods, pipeline architecture, and how data is prepared and delivered across the platform.
Responsibilities
- Design and develop data processing pipelines for structured and unstructured data
- Extract features from unstructured text (e.g., expert call transcripts) and normalize semi-structured data for downstream analysis
- Perform exploratory data analysis and build analytical workflows using Python and Pandas
- Clean, transform, and validate large datasets from a variety of internal and external sources
- Write complex SQL against Postgres (Supabase) to extract, aggregate, and prepare data for analysis
- Build and maintain ETL processes and integrations with external data sources
- Contribute to building and populating the company's centralized data lake, in collaboration with the data engineering team
- Collaborate with in-house AI agents to conduct freeform, exploratory analysis and surface trends across the data lake
- Develop reproducible analyses, notebooks, and internal tooling to support platform decisions
- Partner with full-stack developers where analytical output is surfaced through the product
- Leverage AWS data services (S3, RDS, IAM as they relate to data) to support pipeline and analysis work
- Ensure data quality, lineage, and documentation across pipelines and datasets
- Collaborate with internal stakeholders via GitHub and asynchronous tools
Qualifications
- Strong proficiency in Python for data analysis, particularly with Pandas (required)
- Experience with NumPy and the broader Python data ecosystem (e.g., Jupyter, scikit-learn-adjacent tooling)
- Hands-on experience designing and developing data processing pipelines for structured and unstructured data
- Solid SQL skills, including writing analytical queries against Postgres or similar relational databases
- Familiarity with AWS data services (S3, RDS, IAM as they relate to data access)
- Experience working with large, messy real-world datasets — cleaning, transforming, and validating end-to-end
- Strong problem-solving, debugging, and communication skills
- Ability to own projects end-to-end and work with minimal oversight
- Comfort collaborating with AI agents to conduct freeform, exploratory data analysis
- Moderate proficiency in TypeScript / JavaScript (Node.js) — comfortable using AI coding assistants like Claude Code to work as a cross-functional contributor alongside the full-stack team
Nice to Have
- Machine Learning experience — building, evaluating, or deploying ML models
- Background working with research, data, or publishing platforms
- Understanding of financial services or investment research workflows
- Experience with data warehousing tools, analytics pipelines, or reporting layers
- Experience with workflow orchestration tools (e.g., Airflow, Prefect, Dagster)
- Experience with Vercel deployments and Next.js applications (to collaborate with the app team)