C CRYSS
Download PDF
CRYSS White Paper

CRYSS: A Data Analysis System for High-Throughput Crystallization Screening

CRYSS is an open-source insight generator for partitioned equilibrium datasets, designed to convert complex multicomponent solubility behaviour into mechanistic, decision-ready purification models.

Status V1.0 Live
Scope Industrial & Academic
License MIT

Executive Summary

Purification is often the decisive stage in process development. In multistep synthesis, even a sequence with strong individual yields can still fail economically because every impurity-removal step matters. A ten-step route operating at 90% yield per step delivers only about 35% overall yield, so the real challenge is not simply making the target molecule, but isolating it in a form that is both pure and scalable.

Crystallization is one of the most attractive purification technologies for a solid target because it is simple, robust, and compatible with large-scale manufacturing. Yet the choice of solvent system is rarely obvious. The best solvent is not necessarily the one that dissolves the most material, but the one that gives the sharpest selective partitioning between the target and the impurity profile. In practice, this means identifying conditions that maximize recovery while maintaining the desired purity and processability.

The challenge is compounded by the number of variables involved. Solvent identity, solvent mixtures, temperature, composition, and the possibility of salt formation or derivatization all influence the outcome. Traditional empirical screening quickly becomes unmanageable. CRYSS was developed to close that gap by converting equilibrium-partitioned datasets into mechanistic, decision-ready models of purification performance.

At the core of CRYSS is the concept of Rmax: the maximum physically achievable recovery of the target material under the fitted equilibrium model. Rather than relying on trial-and-error interpretation alone, CRYSS fits the data, compares multiple candidate systems, and ranks them by their predicted purification ceiling. This allows teams to focus experimentation on the most promising solvent systems before committing to costly optimization work.

Rmax = maximum physically achievable recovery of the target material under the fitted equilibrium model

CRYSS Workflow Overview

  1. Equilibrate the candidate mixture with various solvents or solvent mixtures in a closed system to avoid solvent loss. Achieving equilibrium is essential, so thermal cycling, agitation, and sufficient residence time are required.
  2. Determine the mass fraction of the mixture in solution and solid phases at equilibrium. If performed carefully, only calibrated analytics of the solution phase are needed: the initial mass and dissolved mass directly imply the solid mass by difference. Solvent loss must be controlled because evaporation can artificially inflate the apparent dissolved mass.
  3. Record each experiment in a structured dataset, including starting mixture composition Xc, solid-phase composition Xs, solution-phase composition Xl, and experiment or mixture ID.
  4. Fit the dataset in CRYSS, which ranks each experiment by the predicted maximum achievable yield, Rmax, under the selected process conditions.

The workflow therefore converts equilibrium partitioning data into a mechanistic view of process viability. It produces a ranked list of systems worth deeper development and highlights the mixtures most likely to preserve yield while rejecting impurities and by-products.

CRYSS workflow summary diagram
Figure 1: CRYSS enables a highly automated and consistent workflow for equilibrium-based crystallization screening.

Core Principle

In a binary system, if both components are present in the solid at equilibrium, the solution phase must contain the eutectic composition. Knowing that eutectic composition allows direct calculation of the maximum possible recovery of pure material.

In multicomponent systems, by contrast, the solution phase only has the true eutectic composition at Rmax, not throughout the entire recovery window. This distinction is central to the problem: the system is not described by a single constant eutectic throughout the partitioning trajectory, but by a moving composition map that evolves as recovery advances.

The analytical profile across the recovery range from R = 100% to Rmax encapsulates the system-level information required to assess purification viability. It is analogous to diastereomer screening, but the multicomponent case requires much more sophisticated interpretation of composition-space behaviour.

Decision lens CRYSS focuses on the recovery region where the system is mechanistically meaningful, rather than treating a single apparent maximum as sufficient evidence of process quality.
For a binary resolution system: recovery is governed by the eutectic composition of the liquid phase.
For an n-component system: the relevant state space evolves continuously as the system moves from full dissolution toward the optimum recovery condition.

Understanding Multicomponent Eutectics

In an n-component system, the most soluble portion is a eutectic containing all n components. Once dissolved, the next most soluble portion is a eutectic of (n-1) components, and so on until the least soluble portion remains. In the most common scenario, the least soluble component is the desired product, but occasionally a recalcitrant impurity is the least soluble species, leading to purification of the wrong compound. This is a major failure mode in real systems.

Because the starting composition is known, even a single solvent aliquot yields a composition slice at a specific recovery point. Using a mechanistic equilibrium model, CRYSS fits the data and projects forward to estimate the solvent volume and recovery condition corresponding to Rmax.

The power of the method is its scalability: CRYSS can process large datasets automatically, rank systems from high-yielding to poor, and allow rapid triage of large screening arrays without compromising mechanistic interpretation.

In principle, if the fitted parameters are accurate, a single well-constructed experiment is sufficient to estimate Rmax. Additional solvent volumes and replicate measurements strengthen the dataset and improve confidence in the fitted model.

Five-component eutectic visualization
Figure 2: NODES simulation of a five-component system showing how composition slices at different recoveries alter target purity and impurity behaviour.

Performance of CRYSS

Using the NODES application, it is possible to generate arbitrarily large in silico datasets that act as proxies for real experimental data. This enables the selection of meaningful slice points within the critical composition space between R = 100% and Rmax. Each slice point corresponds to a specific solvent volume in a real experiment and provides a different operating state of the system.

These synthetic datasets are internally self-consistent and, critically, the true Rmax value is known for each mixture. This makes them ideal for testing the predictive power and robustness of CRYSS under controlled but realistic conditions.

NODES also supports introduction of realistic experimental error, including analytical error and systematic sampling error caused by cross-contamination between solid and liquor phases. This allows the method to be evaluated against the kind of noise that is encountered in practical laboratory screening.

Single-volume data

Testing CRYSS with a single solvent volume gives a clear relationship between known and predicted Rmax values. Even with minimal data, the scatter plot allows rapid visual separation of robust candidates from weak or non-viable systems.

Single-volume CRYSS performance scatter plot
Figure 3: CRYSS performance using a single solvent volume. The plot compares known NODES Rmax values with CRYSS-predicted Rmax values.

Three-volume data

With three solvent volumes, the predictive relationship improves substantially. The results show near-perfect agreement between true and predicted Rmax values, demonstrating that increasing the number of composition slices materially strengthens the fitted model and reduces uncertainty in the ranking.

Three-volume CRYSS performance scatter plot
Figure 4: CRYSS performance using three solvent volumes. Blue points represent Monte Carlo-only predictions; orange points show the hybrid Monte Carlo + optimizer result.

Summary across all datasets

Across the full test set, the pattern is consistent: error-free datasets perform strongly, while moderate experimental errors become apparent only in the most constrained multi-volume conditions. The important outcome is that CRYSS remains meaningfully robust even when the underlying dataset contains realistic noise.

Across-dataset CRYSS performance summary
Figure 5: Summary of R2 values across all NODES simulation datasets, comparing error-free and 5% noise scenarios.

Technical Approach

CRYSS employs a hybrid strategy combining Monte Carlo simulation with focused optimization. Monte Carlo explores the parameter space broadly, whereas optimization refines the best candidates found during the global search. Together this yields a model that is robust to noise while remaining computationally efficient for screening studies.

  1. Run a broad screen using a single solvent volume per mixture.
  2. Select the top-performing systems using the CRYSS-predicted Rmax.
  3. Perform targeted follow-up experiments on this subset using two or more solvent volumes.

In its present form, CRYSS accepts Excel input files, parses replicates, clusters experiments by starting mixture composition, identifies the relevant R-slices via lever-rule analysis with outlier correction, and solves for Rmax. For real experimental data where the true value is unknown, CRYSS outputs a structured Excel file containing mixture IDs and predicted recovery values.

Rmax is inferred from the composition trajectory between the initial slurry and the equilibrium condition at which the target species yields the maximum physically achievable recovery.

Expected Impact

CRYSS introduces a new capability into crystallization development: the ability to convert small-scale equilibrium partitioning data into mechanistic predictions of maximum achievable yield across multicomponent mixtures. This transforms early-stage purification screening from an empirical exercise into a quantitative, model-driven workflow.

By ranking solvent, solvent-mixture, and salt-forming systems according to predicted purification ceilings, CRYSS enables process chemists to focus effort on the most promising systems instead of optimizing dozens of marginal candidates. This reduces experimental burden, accelerates development timelines, and increases the likelihood of selecting a high-yielding route early in a project.

The method also scales naturally to complex mixtures, salt-forming systems, and derivatization strategies. It integrates cleanly with high-throughput experimentation and automated laboratory workflows, making it especially well suited for future hardware and software integration.

Conclusion

CRYSS V1.0 demonstrates that multicomponent crystallization behaviour can be decoded, modelled, and used to guide purification strategy. The results from NODES benchmarking show that even sparse equilibrium data contain enough information to estimate Rmax with high fidelity, and that hybrid Monte Carlo plus optimization strategies deliver strong predictive power.

The methodology is robust, scalable, and experimentally efficient. A single well-constructed equilibrium experiment can often provide a useful estimate of Rmax, and additional solvent volumes sharpen the prediction further. CRYSS therefore offers a practical route to mechanistic insight without imposing heavy experimental overhead.

Most importantly, CRYSS clarifies the earliest and most consequential stage of crystallization development: selecting which systems deserve attention. By identifying high-yielding solvent and salt-forming conditions before optimization begins, CRYSS ensures that development resources are directed toward the right problems rather than the wrong ones.

Acknowledgements

CRYSS would like to acknowledge the contribution made by Eric Damen, Simone Darphorn-Hooijschuur, Francois Gilardoni, Ben McKay, Lisa McQueen (née Agocs), Erik-Jan Ras, and Goran Verspui.

This draft is formatted for web presentation and intended to support the downloadable PDF version as a companion document.