Astra Trainer
Future Industries

The Bioinformatics Hire Everyone Is Looking For and Nobody Can Find

Aleksandr Mikhailov
Founder, Astra Trainer
Updated
9 min read

Open any bioinformatics vacancy and you find the same list. Strong programming, statistics, pipeline development, cloud infrastructure, and deep understanding of the relevant biology, ideally with domain experience in whatever the company works on.

That person exists. There are not many of them, they are expensive, and they already have a job.

The job description that describes nobody

The reason the vacancy stays open is that it is two careers in one posting.

A competent computational scientist takes years to build. A competent biologist takes years to build. Organisations write a job description requiring both and then treat the resulting difficulty as a market shortage.

A role that requires two full careers is not a hiring problem. It is a job design problem, and it is usually solved on the org chart rather than in the market.

This is the clearest case in the biotechnology domain of the diagnostic that applies across the whole Future Industries section: before concluding the labour market is short, check whether the requirement is fillable as written.

What the direction covers

The scope: reading genomes, sequences and proteins at scale, biological data and computational modelling.

Concretely, four layers.

Sequence analysis. Alignment, assembly, variant calling, annotation. Largely standardised, heavily tooled, and where most of the routine work sits.

Statistics for biological data. Which is its own problem. Biological datasets are wide and shallow, with many more measured features than samples, which breaks the intuitions of people trained on conventional data. Multiple testing, batch effects and confounding are the standard traps.

Structural and protein computation. Modelling structure and interaction, an area that has changed substantially in recent years and continues to move.

Pipelines and infrastructure. Making an analysis run repeatably, at scale, with the versions recorded. Less glamorous and more decisive than anything above it.

The two routes in, and what each one gets wrong

From computing into biology. Arrives able to write reliable code, handle scale, use version control properly and build a pipeline that works. The failure mode is producing a technically correct result that is biologically meaningless, and having no internal alarm that this has happened. Batch effects interpreted as biological signal is the classic example, and it is not a coding error.

From biology into computing. Arrives able to tell whether a result makes sense, which is the thing that cannot be automated. The failure mode is analysis that exists as a folder of scripts nobody else can run, with undocumented manual steps, which works once and cannot be reproduced.

Both are real and both are trainable. The second route is usually shorter, for a reason worth stating plainly: the biological judgement takes longer to build than the engineering discipline, so starting with the person who already has it means training the faster half.

That runs against the instinct to hire a software engineer and teach them some biology, which is the more common approach and the slower one.

Where this sits in the domain

Bioinformatics and computational biology is the fifth of ten directions in Astra Trainer's biotechnology domain, and it is the one most often scoped in combination with others, usually genetics and genomics on the biology side.

For partners training biologists into computational work, the AI, data and computing domain runs eight directions covering computer science, software engineering, data science and analytics, and cloud computing and DevOps, which is where the engineering half comes from. That cross-domain pairing is one of the more common program shapes. You can see the biotechnology directions here.

The layer nobody budgets for

Reproducibility and data engineering. It is boring, it is where the money quietly goes, and it is almost never in the workforce plan.

The symptoms are consistent across organisations.

An analysis cannot be rerun. The person who did it has left, the environment has changed, and the result underpinning a decision cannot be reproduced. In a regulated context this is a serious problem rather than an inconvenience.

Data cannot be found. Sequencing output accumulating without consistent metadata, so datasets from two years ago are effectively lost while still occupying storage that is billed monthly.

Compute costs grow without explanation. Cloud spend that nobody owns, driven by pipelines that were never optimised because nobody was responsible for that.

Every project rebuilds the same thing. No shared tooling, so each scientist writes their own version of the same analysis, slightly differently.

These are all data engineering problems and they are solved by people with software and infrastructure backgrounds rather than by hiring another computational biologist. Organisations that recognise this early spend considerably less over a five-year horizon.

The roles, split properly

The practical move is to stop trying to hire one person and define three.

RoleWhat they doTrained from
Bioinformatics analystRuns established pipelines, interprets output, works with scientists on specific questionsBiologists, laboratory staff, genomics technicians
Bioinformatics engineerBuilds and maintains pipelines, infrastructure, reproducibility and cost controlSoftware engineers, data engineers, IT infrastructure
Computational biologistNovel method development, modelling, research-level workLong route, genuinely scarce, hire where essential

Most organisations need a lot of the first, some of the second and very few of the third. Most job descriptions describe the third and expect them to do all of it.

Splitting the role does two things at once. It makes the positions fillable from populations you already have, and it stops expensive senior people spending their time on pipeline maintenance.

Where regulation touches this. Bioinformatics supporting clinical diagnostics, regulatory submissions or GxP environments carries validation, version control and audit requirements that go well beyond good scientific practice. A pipeline producing clinical results is a regulated system. Training builds both the computational capability and the awareness that this boundary exists, and site-specific validation and qualification remain separate obligations.

Building it rather than buying it

A practical sequence for organisations that have concluded they cannot hire their way out.

Start with the biologists who are already doing analysis badly. Most research groups contain someone who has taught themselves enough scripting to get by. They have the domain judgement and self-taught habits that will not scale. Structured training converts them faster than anyone else available, and they are already familiar with the organisation's data.

Bring in engineering capability separately. From your own software or IT function if one exists. They do not need deep biology to build reliable infrastructure, and pretending they do delays the hire indefinitely.

Treat reproducibility as a requirement from the start. Retrofitting it is considerably more expensive than building it in, and the point at which people notice is usually the point at which it is most expensive.

Keep it continuous. Tools and methods in this field move quickly, so a one-off course dates fast. This is one of the areas where short ongoing learning genuinely outperforms an event, for the same reasons discussed in the hub article on microlearning.

What to take from this

The standard bioinformatics job description asks for two careers and then blames the market.

Both entry routes have a characteristic failure, and the biology-to-computing route is usually shorter because the judgement half takes longer to build than the engineering half.

Reproducibility and data engineering are the unbudgeted layer, and the cost of skipping them arrives as lost data, unrepeatable analyses and cloud bills nobody owns.

Split the role into analyst, engineer and computational biologist. Most organisations need many of the first, some of the second and almost none of the third.

And start with the biologist who has already taught themselves to script. They are the fastest conversion in the building.

Frequently asked questions
Why is bioinformatics so hard to hire for?

Because the typical job description requires a strong computational scientist who is also a working biologist. That is two careers, and the shortage is partly a job design problem rather than a labour market one.

Is it easier to train a biologist to code or a coder to do biology?

Usually the biologist, because biological judgement takes longer to build than engineering discipline. Training the person who already has the slow half means training the faster half.

What gets missed in bioinformatics workforce plans?

Data engineering and reproducibility. The costs arrive later as analyses that cannot be rerun, data that cannot be found, and compute spend nobody owns.

How should the role be structured?

As three roles: analyst, engineer and computational biologist. Most organisations need many analysts, some engineers and very few of the third, and hiring one person to be all three is why the vacancy stays open.

Where does this fit in the domain?

Fifth of ten directions in Astra Trainer's biotechnology domain, frequently paired with genetics and genomics, or with directions from the AI, data and computing domain for the engineering half. You can see them here.

Stop advertising for two careers in one job
Ten directions across biotechnology and the bioeconomy, including bioinformatics and computational biology, and eight more across AI, data and computing for the engineering half. Scoped with your own scientists and mapped onto roles you can actually fill.
Written by Aleksandr Mikhailov
Founder, Astra Trainer · Published · Updated
Continue reading