Bright Vision Technologies

Reinforcement Learning Engineer

United States full-time Senior $100,000 - $150,000
full-time Senior level Technology & IT Salary listed Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

Reinforcement Learning Engineer - Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: Reinforcement Learning Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Key Responsibilities
Design and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments.

Develop, calibrate, and maintain simulation environments suitable for large-scale agent training.

Implement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods.

Engineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints.

Apply offline RL and imitation learning techniques where exploration is costly or unsafe.

Use RLHF, DPO, and related techniques for fine-tuning large language models when relevant.

Build scalable training infrastructure for distributed RL, including efficient experience collection and replay systems.

Optimize training stability and sample efficiency through algorithmic and engineering improvements.

Design rigorous evaluation protocols, including out-of-distribution and adversarial test cases.

Implement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight.

Collaborate with applied scientists and product teams to identify high-value RL use cases.

Monitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users.

Document methodology, design decisions, and operational characteristics for internal stakeholders.

Stay current with RL research and translate promising techniques into production-ready solutions.

Required Qualifications
Master’s or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience.

Six or more years of combined RL research and engineering experience.

Strong proficiency in Python and modern deep learning frameworks.

Hands-on experience with at least one major RL library or in-house RL stack.

Solid understanding of probability, optimization, and the theoretical foundations of RL.

Experience designing and tuning reward functions in non-trivial environments.

Familiarity with simulation environments and large-scale experience collection.

Experience training neural network policies on GPU clusters.

Strong written and verbal communication skills.

Track record of shipping or publishing impactful RL work.

Preferred Qualifications
Experience with RLHF for large language models.

Familiarity with multi-agent RL or hierarchical RL.

Exposure to robotics, control systems, or autonomous driving.

Publications in RL or related research venues.

Open-source contributions to RL libraries or environments.

How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at .

Bright Vision Technologies is an Equal Opportunity Employer.
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Originally posted on Himalayas

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Senior
Technology & IT Human Resources Python Remote Collaboration Reinforcement Learning Senior level

Likely questions

  1. Tell us about work you have done that is close to the Reinforcement Learning Engineer role.
  2. How would you approach your first 30 days at Bright Vision Technologies?
  3. Which of Human Resources, Python and Remote Collaboration have you used recently, and what did it help you achieve?
  4. How have you led people, improved a process, or made a hard decision in a previous role?
  5. How do you handle busy days, changing priorities, or pressure at work?

Prepare before the call

  • A recent example that proves your experience with Human Resources, Python and Remote Collaboration.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the Reinforcement Learning Engineer role because I can bring practical experience in Human Resources, Python and Remote Collaboration, learn the team quickly, and contribute to the outcomes Bright Vision Technologies needs from this hire.

Related jobs.

More roles from this company or category.