flavormancer

mcp
Security Audit
Warn
Health Warn
  • License Ò€” License: Apache-2.0
  • Description Ò€” Repository has a description
  • Active repo Ò€” Last push 0 days ago
  • Low visibility Ò€” Only 8 GitHub stars
Code Pass
  • Code scan Ò€” Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions Ò€” No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Predict taste & aroma from a molecule's chemical structure β€” on-prem flavor prediction for food, beverage & fragrance R&D. πŸ§ͺ

README.md

Flavormancer

Flavormancer β€” taste & aroma prediction from chemical structure

On-prem flavor prediction from chemical structure β€” taste and aroma, physicochemical
behavior, formulation notes, and safety flags
for any molecule, plus substitution search.
Enter a common or IUPAC name (or a SMILES) and get a single, honest flavor read,
running entirely on hardware you own.

A full, confidence-tagged molecule flavor read

A single, confidence-tagged flavor read β€” taste, aroma, behavior, safety, and drop-in substitutions.

Flavor-space map in 3D on MW Γ— logP Γ— TPSA axes, colored by taste and aroma

The interactive flavor-space map in 3D on real MW Γ— logP Γ— TPSA axes, colored by taste & aroma β€” every one of the 164 aroma + 6 taste classes labelled. 187 trained heads in all: 6 taste + 164 aroma + 5 mouthfeel + 12 safety.

8,847 unique molecules across the open datasets Β· taste + aroma + mouthfeel prediction from
structure Β· a flavor library (start from a flavor β†’ its character-impact molecule) and
flavor designer (pick your notes β†’ best food-safe molecules + drop-in swaps) Β· an
interactive 2D/3D flavor-space map Β· 2D & 3D structure views. All on
commercial-clean public data.

Per set (unique molecules): taste training 3,845 Β· aroma training 2,394 Β· odor
corpus 2,255 Β· documented taste 676 Β· mouthfeel training 2,534 Β· Tox21 (safety)
7,823 Β· GRAS reference 2,781 Β· sweetness intensity 316 Β· character-impact aroma
supplement 602 associations (open-gov-sourced). Every one of the 8,847 is enriched with
names + measured properties from public-domain PubChem.

How the universe grows: the flavor-space map and enrichment table show 8,847 unique
structures
(deduped by connectivity skeleton), expanded by ingesting the full EU/GB
flavourings Union List (~2,200 authorised, Open Government Licence v3)
so the browse-able
universe is food-forward. Meaningfully-distinct stereoisomers (e.g. R- vs S-limonene) that carry
their own documented odor/taste are surfaced as first-class rows; nothing is dropped.

This is the commercial edition β€” Apache-2.0 and commercial-clean: every
model (taste and aroma) trains only on permissively-licensed / public-domain
data, so it's free to use, sell, and run behind your firewall. The commercial-clean
aroma model ships here as presence/absence; a higher-fidelity academic
edition
adds the scored-intensity model from research odor data (NonCommercial
terms) β€” see Two editions.

Status: the Python training pipeline and prediction core are built and
tested; the .NET serving layer, React workbench, and packaging are in progress.
See the roadmap β€” M0–M1 landed, M2–M6 next.


What it does

Given one molecule, Flavormancer returns a structured profile where every value
is tagged by how it was derived
, so nothing reads as more certain than its source:

  • Taste β€” six trained heads: sweet / bitter / umami / sour / salty /
    tasteless (RandomForests on fingerprint + physicochemical features), plus a
    sweetness-intensity regressor. Sour and salty also keep a transparent chemistry
    rule (acid group / alkali-salt) as a deterministic cross-check alongside the model.
  • Aroma β€” 164 odor-descriptor heads (citrus, floral, minty, almond, fatty,
    petroleum, earthy, medicinal, sulfurous, camphor, fruity, fishy, garlic, ethereal,
    ammoniacal, pungent, pine, rose, rancid, alcoholic, woody, green, grassy, putrid)
    trained on public-domain HSDB odor text + curated character-impact facts, surfaced
    for any molecule, each with its held-out CV-AUROC (0.71–0.98). Presence/absence,
    honestly β€” intensity is the "comes with your data"
    upgrade (no public intensity data exists). Documented odor + detection thresholds are
    shown where cited.
  • Behavior β€” logP, molecular weight, TPSA, H-bonding, ring/atom counts
    (computed); water solubility (ESOL estimate); volatility tier and pKa ranges
    (qualitative); measured boiling point / vapor pressure when a property table is
    loaded (lookup β€” structure-based BP was evaluated and declined as too inaccurate).
  • Mouthfeel β€” five trained chemesthesis heads (cooling / pungent / warming /
    astringent / tingling): the trigeminal sensation, trained on curated public-domain
    agents (menthol & WS-coolants, capsaicinoids, tannins, Sichuan-pepper sanshools), each
    with its held-out CV-AUROC. Distinct from the same-named aroma notes β€” the sensation, not
    the smell (menthol feels cool; WS-23 cools with almost no odour).
  • Stability β€” oxidation / hydrolysis / photo watch-flags.
  • Safety (defensive, caution-only) β€” a disclaimer + scope on every result,
    twelve Tox21 in-vitro assay heads (nuclear-receptor + stress-response, each with its
    CV-AUROC) surfaced as indicative review flags, structural tox-alert screening, a
    preliminary TTC/Cramer concern tier, an optional GRAS cross-reference, and EU
    declarable-allergen labeling. It flags for review; it never clears a compound for use.
  • Formulation β€” a documented dangerous-mixture screen (benzene, nitrosamine,
    acrylamide, ethyl carbamate, furan, and more) and an OAV dosing-balance analysis
    that flags the component about to overpower a blend (quantitative when threshold
    tables are loaded).
  • Substitutes & structural neighbors β€” two nearest-neighbor searches over the whole
    molecule universe for reformulation and cost-down: substitutes rank by taste + aroma
    profile
    match (a molecule that tastes and smells like the target β€” e.g. ethyl vanillin for
    vanillin β€” regardless of structure), while structural neighbors rank by Tanimoto/Morgan
    structure similarity (the look-alikes).
  • Flavor Studio β€” one hub to pick any mix of everyday flavors (banana,
    saffron, pumpkin, bubble gum…) and notes (citrus, floral…) β†’ ranked food-safe
    molecules + drop-in swaps. A flavor is a set of notes, so they live in one picker.
  • Flavor-space map β€” every molecule embedded 2D/3D: a similarity layout
    (UMAP over fingerprints) and an interpretable property-axes layout (MW Γ— logP Γ—
    TPSA), colored by taste, aroma, or both.
  • Chirality explorer β€” enumerates every stereoisomer (R/S and E/Z), surfacing
    the ones with distinct documented odor/taste (R- vs S-carvone) as their own reads.
  • Master enrichment table β€” sortable/searchable grid of the whole universe with a
    "why it matters" note on every chemistry column.
  • External references β€” per-molecule deep links to PubChem (public domain) and
    NIST WebBook for spectra (IR/MS/NMR) and GC retention indices, with an availability
    flag showing what each source has. We link, never rehost.

On aroma intensity (the one gated piece). The aroma heads above are
presence/absence β€” which notes apply, from public-domain data. Scored
intensity
(how strong) is the marquee upgrade that comes with a customer's own
panel data or a commercially-licensed set (Leffingwell PMP 2001); no
public-domain intensity data exists. That's the only aroma capability still gated.

What aroma training needs from you: molecules (SMILES, or GC-MS to identify
the compounds in your products) paired with your panel's expert odor descriptors
(e.g. green / fruity / woody, ideally with intensity). GC-MS identifies the
molecules; the sensory labels are what the model learns β€” GC-MS alone isn't enough.

For exactly what data unlocks each further capability (aroma, quantitative dosing,
retention index) and the formats we accept from a client, see
docs/DATA-REQUIREMENTS.md.

Scope: Flavormancer predicts flavor properties only. It is not a safety,
toxicity, GRAS, regulatory, or stability determination. A prediction is never a
clearance to consume.

Confidence tiers

Every output is labeled by how it was produced: computed (exact from
structure), trained (ML on open data), rule (deterministic structural
rule), estimate (published QSPR with known error), lookup (from a loaded
reference table), and qualitative (a class/flag, not a number). Full map in
docs/CAPABILITIES.md.

How does it all actually work? A plain-English tour of the method behind every
feature β€” fingerprints, the taste/aroma models, the map, chirality, the mixture
reactions β€” is in docs/HOW-IT-WORKS.md.

Why it's built this way

The value isn't any single number β€” those you can look up. It's four design choices:

  • Prediction, not lookup. The trained models read a molecule from its structure,
    so they answer for a novel or unmeasured compound that's in no database β€” not just
    for known ones.
  • On-premise, behind your firewall. Proprietary candidate structures never leave
    your network
    . You can screen confidential molecules without exposing your direction
    to any external service β€” something no public web tool can offer.
  • One integrated read. Taste, behaviour, safety, and substitution in a single
    confidence-tagged screen, instead of stitching together half a dozen databases and
    manual checks per molecule.
  • It sharpens on your own data. The open-data models are the floor. Trained on a
    user's own formulation and sensory data β€” on the same on-prem box β€” they become
    specific to that user's products, which no public dataset can be.

Two editions

Flavormancer ships as two editions of one method:

Flavormancer (this repo) Flavormancer Research
Edition Commercial Academic / open-source (coming soon)
License Apache-2.0 open-source, research / NonCommercial
Data commercial-clean open data only adds research odor datasets with NonCommercial terms
Aroma 164 presence/absence descriptor heads ship (public-domain HSDB); scored intensity is trained on your data or a licensed set (PMP 2001) full open model incl. intensity (research odor data)
Use free to use, sell, run on-prem research, teaching, advancing the method

The split is deliberate. The richest aroma data is licensed for research only, so
it can't ship in a product you sell β€” keeping it out is exactly what makes this
edition clean to use and sell, and the academic edition is where that fuller model
lives.

Architecture

Python trains the models offline; a .NET application serves them at runtime β€”
nothing at runtime depends on Python.

data sources ─► Python training (build-time) ─► ONNX (taste) ──┐
                RDKit Β· scikit-learn                            β”‚
                                                                β–Ό
React workbench ◄─ JSON API ◄─ ASP.NET Core + ONNX Runtime + Postgres/pgvector
                                              β”‚
                                              β–Ό
                              Docker Compose on a single on-prem box
Layer Technology
Model training (build-time) Python Β· RDKit Β· scikit-learn Β· skl2onnx
Model handoff ONNX
App / API ASP.NET Core (C#)
ML serving ONNX Runtime, in-process in .NET
Frontend React
Database PostgreSQL + pgvector
Deploy Linux + Docker Compose (single box)

Repository layout

training/        Python β€” dataset build + model training (build-time)
api/             ASP.NET Core β€” app, auth, endpoints, ONNX serving
frontend/        React β€” the workbench UI
infra/           Dockerfiles, docker-compose.yml, deploy
docs/            architecture, capabilities, and design docs
tests/           pytest suite for the prediction core

Getting started

See training/SETUP.md for the clean-machine setup. Datasets and
trained models are not committed β€” the training scripts pull their sources and
.gitignore keeps artifacts out of the repo.

Team

Built by Echelon Technology Solutions β€” a small team where
everyone is cross-training toward full-stack.

Role
Austin Β· @rvnminers-A-and-N Founder & lead β€” training pipeline, backend, architecture, models, and roadmap; mentors the team across their areas
Jamie Front-end lead β€” owns the UI/UX
Aaron Β· @Jabbacado Full-stack (in training) β€” backend, infrastructure, AI, UI/UX
Ty Β· @Frostfire0101 Full-stack (in training) β€” backend, infrastructure, networking, AI, UI/UX

Austin β€” B.S. Chemistry (ACS-certified) & B.S. Physics, magna cum laude, Western
Kentucky University (minors in Mathematics & Astronomy), plus graduate coursework in a
computational-chemistry M.S. program. Background in molecular-dynamics simulation β€”
protein conformational states and SOβ‚‚ at the air–water interface β€” on the WKU HPC
cluster: the structure-to-property grounding behind Flavormancer.

Contributing

See CONTRIBUTING.md for branching, commit conventions, and the
review workflow. Commits are signed and signed off (DCO + CLA); PRs are
small, linked to an issue, and squash-merged after review.

License

Licensed under the Apache License 2.0 β€” code license only; datasets and
any pretrained models carry their own licenses, noted where used. The academic
edition is a separate repository under its own (NonCommercial) terms.

Reviews (0)

No results found