AI Code Evaluation Specialist - Freelance

Department Icon Data Science Analytics & Machine Learning
149+ Applicants
Posted: 5 months ago
3-5 years
Delhi, Delhi
work from office

Posted: 5 months ago
|
Applicants: 149+
Job Description
Similar Jobs
Please verify your account first! Send OTP

Please click on the Apply to verify the status of jobs posted more than 15 days ago, as they may have expired. Similar Jobs

Job Description

The Opportunity

Evaluate cutting-edge AI-generated code through real-world engineering challenges. Work with major open-source projects — data orchestration engines, ML/AI model libraries, LLM tooling frameworks, and workflow automation platforms. Your structured, evidence-based evaluations help train and benchmark the next generation of AI coding assistants.

About Biz-Tech Analytics

Biz-Tech Analytics (BTA) is a data services and AI solutions company specializing in high-quality labeled datasets, training data, and AI model evaluation. We partner with leading AI research teams and enterprises to deliver the data infrastructure that powers responsible AI development.

What Youll Do
  1. Analyse Real Engineering Challenges: Work with complex engineering tasks sourced from major open-source projects (e.g., implementing new schedulers, adding model architectures, building tool integrations).
  2. Write Engineering Prompts: Craft clear, self-contained prompts that describe engineering tasks for AI coding models to attempt.
  3. Set Up Development Environments: Prepare local development setups using Git, VS Code, tmux, and Python. Clone repos, create branches, and configure dependencies.
  4. Run Evaluation Workflows: Use provided tooling to generate AI model outputs on engineering tasks and capture the resulting code changes.
  5. Review Code Changes: Thoroughly examine every file change in diffs. Validate correctness, code quality, adherence to best practices, and alignment with original requirements.
  6. Complete Structured Evaluations: Write 500–800 word evaluations assessing AI-generated code across multiple dimensions including correctness, code quality, and engineering judgment. Provide evidence-backed ratings using structured rubrics.

You Must Have
  • 3+ years of professional software development experience
  • Strong proficiency in Python AND at least one of: JavaScript/TypeScript, Go, Rust, Java, or C++
  • Experience reading and reasoning about large, unfamiliar codebases — including tracing cross-file dependencies across 10-50+ file changes
  • Solid understanding of software testing: unit/integration tests, test frameworks (pytest, Jest, Mocha), and ability to assess test completeness and coverage gaps
  • Git fluency: branching, PRs, diffs, cherry-picks, and merge conflict resolution
  • Ability to set up and troubleshoot development environments (Docker, tmux, VS Code, terminal/CLI workflows)
  • Strong written English: ability to compose detailed, structured evaluation reports (500+ words, clear logic, professional tone)
  • Code review experience: ability to read code critically, trace architectural impact across files, assess backward compatibility, and identify correctness, performance, and maintainability issues

Bonus Points
  • Looking to get Placed? Try our Placement Guarantee Plan

    Contributions to open-source projects (data orchestration, ML/AI libraries, LLM tooling, workflow automation, or similar)
  • Experience with ML/AI frameworks (PyTorch, TensorFlow, JAX) — strongly preferred as 40% of tasks involve ML-related codebases
  • Familiarity with AI code generation tools (GitHub Copilot, Claude, ChatGPT)
  • Background in QA/testing roles, formal code review, or experience across diverse domains (backend APIs, data pipelines, ML systems, developer tooling)

Please Dont Apply If
  • Youve only done data engineering/ETL without application development
  • You struggle to read and reason about code you didnt write
  • Youre uncomfortable with terminal/CLI workflows
  • Your written English isnt strong enough for structured 500+ word evaluations
  • Youre looking for a pure coding/development role (this is evaluation-focused)

Skills

PythonEtlAnalyticsAiMl

If a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.

Important dates & deadlines?

Application Deadline

28 Jun 26, 03:52 PM IST

Similar Jobs

View All
Loading...
Bag Logo
Jobaaj
Don't Miss out any Updates

Subscribe now for the latest job alerts
and never miss an update

Job Alert
Google hiring for Specific Roles Apply Now!
1 min ago
New Opportunity
Amazon is hiring freshers Apply Now!
5 min ago
Featured Jobs
Microsoft opening 50+ positions Apply Now!
10 min ago

AI Code Evaluation Specialist - Freelance

Share with