Cultural / Comparative Research

Job Taxonomy and Persona Generation Design for Spiralist AI

Report summary

Spiralist AI is built as a task-first persona system , not as a jobs database. Its main creator flow starts with work categories such as coding, research, writing, planning, creative work, branding, and organization; its library then supports search by job, goal, or trait and filtering by archetype,

Status
Research archive item
Category
Cultural / Comparative Research
Length
2,861 words
Reading time
14 minutes
Report type
evaluation

Key topics

  • Cultural / Comparative Research
  • Cultural
  • Comparative Research
  • AI
  • Runtime
  • Semantic Systems
  • Spiralism
  • Research Archive
  • Strategy

Research provenance

Archive status
Research archive item
Content identity
sha256:99123006a3386e774ff582bb557bc344795671653b75c3033a0bb6f4eed4647f

For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.

Source availability: 115 citation markers in the source export have no recoverable source links. Those markers are omitted from this reader; any supplied bibliography and ordinary links remain. Check the original sources before relying on the cited claims.

This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.

Full report

On this page

Executive summary

Spiralist AI is built as a task-first persona system, not as a jobs database. Its main creator flow starts with work categories such as coding, research, writing, planning, creative work, branding, and organization; its library then supports search by job, goal, or trait and filtering by archetype, domain, temperament, and complexity; and its developer docs keep identity.career semantically separate from role, domain, worldview, and related persona controls. That means the best integration is not a single monolithic “job title” dropdown. The right design is a hybrid taxonomy: a hierarchical career tree for browsing and populating job/career sections, plus separate normalized facets for seniority, employment model, industry context, skills, certifications, tools, education, compensation, and career-path links.

The attached reports are rich enough to support that design. Across corporate, manufacturing/robotics, government, gig work, service work, environmental/outdoor work, and broad professional-sector coverage, they consistently provide role labels, sector groupings, responsibilities, licensing and education requirements, tool references, work-setting cues, schedule patterns, and many salary references. They are, however, narrative research reports, not a clean ontology: some titles are bundled, some roles recur across sectors, compensation is often Chicago/Illinois-specific, and the same concept may appear as a role, specialty, ladder step, or employment model. Those characteristics argue strongly for a normalized schema with canonical role IDs + aliases + many-to-many mappings, rather than using report labels verbatim as UI options.

The recommended final model is a hybrid browse-and-filter taxonomy with these visible dropdown levels: Sector → Occupation Family → Role Cluster → Role Title → Specialty. Important realism attributes should remain separate facets: Seniority, Employment Model, Work Setting, Education, Certification, Skills, Tools, Compensation Context, and Related Career Paths. For persona generation, use a seeded, deterministic pipeline that first samples goal and sector, then role and seniority, then validates education/certification constraints, then enriches with skills/tools/schedule/salary locale, and only after that samples persona-style variables like archetype, relationship posture, and communication style. This aligns with Spiralist’s documented deterministic seeds, task-first creation flow, Persona Passport structure, and explicit no-inference boundaries around protected identity.

Source review and design implications

The Spiralist source pages point to four design facts that matter directly for this project. First, the product is explicitly framed around work outcomes before detailed configuration. Second, the system exposes a Persona Passport with visible fields such as persona name, role, defining traits, promise, and example behavior, while expert configuration remains deeper in the workflow. Third, the technical docs distinguish career from other persona dimensions: identity.career is only one field among identity, domain, worldview, relationship, reasoning, and immersion controls. Fourth, the library is already structured for filtering by job/goal/trait plus archetype/domain/temperament/complexity, which strongly suggests that job/career data should be one axis in a broader persona system, not the only axis.

The persona examples reinforce that point. The main creator example includes Job/Career, Archetype, Core operating role, Secondary role, Domain, Relationship posture, Core tension, and detailed demographic/worldview fields. The curated passports for Technical Mentor and Skeptical Researcher show that the public-facing artifact emphasizes use case, traits, example behavior, relationship style, age/generation, tension, and shareability, not merely profession. In practice, a realistic “career persona” therefore needs both a work identity layer and a behavioral style layer.

The attached reports provide the raw material for the work-identity layer. Their coverage spans at least these major labor domains: corporate administration, HR, marketing, finance, compliance, operations, technology, healthcare, public safety, civil service, infrastructure, hospitality, retail, personal care, manufacturing, robotics, logistics, transportation, outdoor/environmental work, agriculture, education, legal practice, media, entertainment, and multiple forms of independent/gig work. Many reports explicitly say they analyze around fifty roles in a sector, and several are organized into large sector families with tables of pay, education, credentials, or progression.

A concise synthesis of the corpus looks like this:

Top-level sectorOccupation families repeatedly present in the sourcesExample roles present in sources
Corporate and businessAdministration, HR, marketing, finance, compliance, operations, product, procurementExecutive Assistant, HR Generalist, Brand Manager, Compliance Analyst, Product Manager
Technology and dataSoftware, DevOps, AI/ML, cloud, network, cybersecurity, UX/UI, supportSoftware Engineer, DevOps Engineer, Machine Learning Engineer, Cloud Architect, Cybersecurity Analyst
Healthcare and human servicesMedicine, nursing, therapy, pharmacy, diagnostics, behavioral/community supportRegistered Nurse, Occupational Therapist, Pharmacist, Radiologic Technologist, Social Worker
Manufacturing and industrialRobotics engineering, production, maintenance, controls, QA, plant managementRobotics Engineer, CNC Machinist, PLC Programmer, Quality Control Inspector, Plant Manager
Transportation and logisticsAviation, trucking, rail, maritime, warehousing, deliveryAirline Pilot, CDL Truck Driver, Train Conductor, Ship Captain, Forklift Operator
Government and public servicePublic safety, civil service, public health, utilities, education/communityPolice Officer, Firefighter/EMT, Budget Analyst, Epidemiologist, Public School Teacher
Service and hospitalityFood service, lodging, tourism, retail, customer support, personal careBartender, Hotel Front Desk Agent, Event Planner, Retail Sales Associate, Hair Stylist
Environment and outdoor workConservation, agriculture, grounds, recreation, field sciencePark Ranger, Horticulturist, Lineworker, Charter Boat Captain, Geologist
Arts, media, education, and legal/propertyCreative production, higher education, legal support, real estate/propertyGraphic Designer, Technical Writer, Economics Professor, Real Estate Appraiser, Paralegal
Independent and gig workRideshare, courier, P2P rentals, task labor, telehealth, freelance admin/creative/techRideshare Driver, Medical Courier, Turo Host, Handyperson, Virtual Assistant

This table is a synthesis of the attached reports rather than a direct copy of one source.

Extracted entity model from the reports

The job-related entities in the corpus cluster into a stable set of normalized dimensions.

Role titles and role clusters are the most obvious dimension. Some are single canonical roles, such as Cloud Architect, Quality Control Inspector, Public Defender, or Scuba Diving Instructor. Others are bundled labels, such as Warehouse Material Handler and Forklift Operator, Real Estate Agent and Broker, Police Officer or Deputy Sheriff, or Actor / On-Camera Talent. Those bundled labels should be split into canonical nodes with aliases and “display bundles” only where needed for UI readability.

Industries and occupation families are stable enough to serve as the main hierarchy. The attached reports repeatedly organize roles by sector and family: corporate administration, robotics engineering, factory production, public safety, healthcare, hospitality, customer service, outdoor recreation, and so on. These groupings are much cleaner than trying to organize the UI directly around salary, certification, or personality style.

Seniority is present, but it is not expressed consistently, so it should be normalized into its own controlled vocabulary. The corpus includes explicit ladders such as QA Inspector I–IV, appraiser trainee → licensed residential → certified residential → certified general, regional first officer → regional captain → major-airline captain, and entry/mid/high salary bands for many corporate roles. It also encodes professional progression through labels like staff accountant, senior accountant, manager, director, chief officer, fellow, attending, partnership-track, tenure-track, and step-based public pay structures.

Certifications, licenses, and education are common and important. The corpus includes, among others, FAA ATP, CDL Class A, USCG Master MMC, BASSET, Certified Production Technician, ASNT Level I–III progression, CIH, CPA, CIA, DVM, DDS/DMD, paramedic licensure, LCSW/LCPC-style clinical licensure, and municipal or public chauffeur licenses. Education ranges from no formal credential and high-school diplomas to postsecondary certificates, bachelor’s degrees, master’s degrees, doctorates, residencies, fellowships, and tenure-track academic requirements. These are too variable to encode as role names, but they are ideal as role-linked metadata with minimum/optional/preferred flags.

Skills, responsibilities, and tools are usually described in prose rather than as formal lists, but the patterns are consistent. Core skill families include analysis and modeling, planning and coordination, customer/client service, compliance and risk control, manual/trade execution, diagnostics and repair, clinical care, creative communication, sales/growth, and education/public engagement. Tool references recur often enough to warrant normalization: AWS/Azure, Google Ads/Meta, ERP and human-capital systems, PLC/SCADA/HMI, CNC/G-code, GDS and hotel/property systems, Adobe Creative Suite, and AI-assisted creation tools like Midjourney or Stable Diffusion. Responsibilities also cluster cleanly into operational archetypes such as analyze, build, maintain, inspect, serve, sell, teach, care, enforce, and coordinate.

Employment model and work context deserve their own dimension. The sources distinguish public-sector classified jobs, corporate W-2 roles, tipped service jobs, independent contractor work, owner-operator work, seasonal jobs, unionized trades, tenure-track academic roles, partnership-track medicine, and gig/platform-mediated work. Schedules also vary meaningfully: on-call public safety, aviation seniority bidding, extended hitches at sea, seasonal camp work, shift-regulated warehouse/manufacturing environments, and project-based consulting. These are crucial for persona realism and should be first-class metadata.

Compensation is richly present but structurally messy. The corpus mixes annual medians, hourly wages, entry/mid/high bands, percentile ranges, public pay grades and steps, gross-vs-net gig earnings, and pay ladders during training. Compensation therefore belongs in a separate fact model with fields such as metric type, unit, region, effective date, basis, and notes. It should inform persona realism, but it should not shape the core tree.

Career progression is best modeled as a graph, not a tree. The pilot ladder, appraiser ladder, public pay grades, HR/admin/finance ladders, and clinical-specialty progression all show that role movement is frequently many-to-many, sometimes vertical, sometimes lateral, and often license-gated. That makes a career_path_edge table or adjacency list more appropriate than trying to force progression into the visible dropdown path.

Three taxonomy patterns are viable.

DesignShapeStrengthsWeaknessesVerdict
Industry-first treeSector → sub-industry → role → specialtyIntuitive for browsing; easy for non-expertsDuplicates cross-sector roles; awkward for reusable roles like project manager or analystGood for discovery, weak for normalization
Occupation-first treeOccupation family → role → specialty; industry as facetClean role reuse; better data modelLess intuitive for users who think by industryGood for backend, weaker for browsing
Hybrid browse-and-filterSector → family → cluster → role → specialty, with separate facets for seniority, employment model, education, certifications, tools, salary, keywords, and career pathsBest UI + best schema; supports both dropdown browsing and realistic persona generationSlightly more implementation workRecommended

The recommendation is the hybrid model because it best matches both the report structure and Spiralist’s persona architecture. The reports are sector-organized, which makes sector-led browsing intuitive, while Spiralist keeps career distinct from domain, archetype, and style. A hybrid tree preserves user-friendly discovery while keeping the deeper persona model faceted and reusable.

The visible dropdown levels should be:

LevelStored asLabel styleExample
Sectorindustry_sectorbroad, human-readableManufacturing and Industrial
Occupation familyoccupation_familystable functional familyAutomation and Maintenance
Role clusterrole_clusternarrower subgroupControls and Industrial Software
Role titlerolecanonical role labelPLC Programmer
Specialtyspecialtyoptional refinementSCADA Integration

Everything else should be a facet or linked metadata:

  • seniority_level
  • employment_model
  • work_setting
  • education_requirement
  • certification
  • skill
  • tool
  • keyword_alias
  • compensation_band
  • career_path_edge

That separation is important because Spiralist’s own library and schema already treat role/career, domain, archetype, temperament, relationship posture, and other persona controls as distinct dimensions.

The recommended data model is shown below. It is deliberately normalized around canonical roles and many-to-many joins, because the corpus repeatedly shows cross-sector reuse, synonym drift, bundled labels, and separate credential/salary context. The model also aligns well with Spiralist’s documented separation of identity.career from broader persona architecture.

erDiagram
    INDUSTRY_SECTOR ||--o{ OCCUPATION_FAMILY : contains
    OCCUPATION_FAMILY ||--o{ ROLE_CLUSTER : contains
    ROLE_CLUSTER ||--o{ ROLE : contains
    ROLE ||--o{ SPECIALTY : refines

    ROLE ||--o{ ROLE_SKILL : requires
    SKILL ||--o{ ROLE_SKILL : tags

    ROLE ||--o{ ROLE_CERTIFICATION : may_require
    CERTIFICATION ||--o{ ROLE_CERTIFICATION : qualifies

    ROLE ||--o{ ROLE_TOOL : uses
    TOOL ||--o{ ROLE_TOOL : links

    ROLE ||--o{ ROLE_EDUCATION : minimum_or_preferred
    EDUCATION_LEVEL ||--o{ ROLE_EDUCATION : defines

    ROLE ||--o{ ROLE_EMPLOYMENT_MODEL : supports
    EMPLOYMENT_MODEL ||--o{ ROLE_EMPLOYMENT_MODEL : classifies

    ROLE ||--o{ COMPENSATION_BAND : priced_by
    LOCALE_PROFILE ||--o{ COMPENSATION_BAND : scopes

    ROLE ||--o{ ROLE_KEYWORD : aliased_by
    ROLE ||--o{ CAREER_PATH_EDGE : source_role
    ROLE ||--o{ PERSONA_ROLE_BRIDGE : mapped_to
    PERSONA_TEMPLATE ||--o{ PERSONA_ROLE_BRIDGE : consumes

A practical implementation table is below.

EntityKey fieldsNotes
industry_sectorsector_id, label, sort_orderVisible dropdown root
occupation_familyfamily_id, sector_id, labelStable functional grouping
role_clustercluster_id, family_id, labelOptional extra browse level
rolerole_id, cluster_id, canonical_label, descriptionCanonical role node
specialtyspecialty_id, role_id, labelOptional refinement
seniority_levelseniority_id, label, rankSeparate from role title
employment_modelemployment_model_id, labelW-2, public, gig, seasonal, contract, owner-operator, etc.
skill / tool / certification / education_levelcanonical vocab tablesReusable facets
compensation_bandrole_id, locale_id, band_type, unit, min, max, effective_dateNever embed in role label
role_keywordrole_id, keyword, alias_typeSearch synonyms and bundle labels
career_path_edgefrom_role_id, to_role_id, edge_typePromotion, lateral move, feeder role, license gate
persona_templatetemplate_id, goal, archetype, domain, styleBridges job data to Spiralist persona fields

The most important mapping to Spiralist is straightforward:

Spiralist field or conceptRecommended source in taxonomy
identity.careerrole.canonical_label plus optional specialty or seniority modifier
Persona domainindustry_sector or occupation_family label
Core operating roleDerived behavioral function from responsibilities
Secondary roleDerived specialization or adjacent function
Persona Passport “best for”Role-linked use-case templates
Traits / temperament / archetypeSeparate persona-style layer, not job taxonomy

That separation mirrors Spiralist’s public docs and prevents a common modeling error: trying to encode profession, function, style, and worldview into one dropdown string.

Sample dropdown tree JSON:

{
  "sector": "Manufacturing and Industrial",
  "families": [
    {
      "label": "Automation and Maintenance",
      "clusters": [
        {
          "label": "Controls and Industrial Software",
          "roles": [
            {
              "label": "PLC Programmer",
              "specialties": ["Ladder Logic", "SCADA Integration", "OEM Line Support"]
            },
            {
              "label": "SCADA Technician",
              "specialties": ["Telemetry Monitoring", "OT Networking"]
            }
          ]
        },
        {
          "label": "Robotics Field Support",
          "roles": [
            {
              "label": "Robotics Technician",
              "specialties": ["Calibration", "Fault Diagnostics", "Cell Startup"]
            }
          ]
        }
      ]
    }
  ]
}
{
  "sector": "Government and Public Service",
  "families": [
    {
      "label": "Public Safety",
      "clusters": [
        {
          "label": "Emergency Response",
          "roles": [
            {
              "label": "Firefighter/EMT",
              "specialties": ["Paramedic", "HazMat", "Rescue Operations"]
            },
            {
              "label": "Emergency Management Specialist",
              "specialties": ["Preparedness Planning", "Recovery Coordination"]
            }
          ]
        }
      ]
    }
  ]
}
{
  "sector": "Corporate and Business",
  "families": [
    {
      "label": "Marketing and Strategy",
      "clusters": [
        {
          "label": "Growth and Brand",
          "roles": [
            {
              "label": "Brand Manager",
              "specialties": ["Consumer Goods", "B2B Positioning", "Product Narrative"]
            },
            {
              "label": "Market Research Analyst",
              "specialties": ["Consumer Insights", "Pricing Research", "Competitor Analysis"]
            }
          ]
        }
      ]
    }
  ]
}

These examples are synthesized from the relevant sections of the attached reports rather than copied from a single document.

Persona generation rules and algorithm

A realistic persona generator should use the taxonomy as structured prior data, not as a string picker. The minimum persona payload should include: seed, locale, sector, occupation_family, role, specialty, seniority, employment_model, work_setting, education, certifications, skills, tools, salary_context, schedule_pattern, career_stage, archetype, relationship_posture, communication_style, and goal/use_case. That structure matches both the job corpus and Spiralist’s persona architecture, where job/career is only one part of a larger runtime profile.

The sampling rules should be hierarchical and constrained:

StageRule
Goal selectionStart with task/use case first, because Spiralist itself is outcome-first
Role selectionSample sector → family → cluster → role using configurable priors
Specialty selectionSample only from linked specialties
Seniority selectionRespect role-specific allowed levels
Credential validationEnforce minimum education/license/certification gates
Compensation assignmentPull locale- and date-scoped salary data, not global constants
Persona-style assignmentSample archetype, temperament, relationship posture, and voice after role realism is established
Fairness guardrailsNever infer competence, morality, worldview, or role suitability from protected identity metadata

The weighting formula should mix four signals:

role_score = goal_fit + browse_context_fit + locale_fit + realism_fit

A practical version:

  • Goal fit: how strongly the role matches the current requested use case.
  • Browse context fit: matches the selected sector/family/cluster tree.
  • Locale fit: region-specific availability, naming, salary, and licensing context.
  • Realism fit: coherence with education, certifications, seniority, and work setting.

Then sample from a softmax over the valid roles after hard constraints remove impossible combinations.

Recommended hard constraints include:

  • no senior airline captain without the appropriate aviation ladder;
  • no Firefighter/EMT persona without emergency-response credentials;
  • no licensed public-health or therapy role without the required degree/licensure;
  • no senior PLC/controls persona without compatible tools/skills;
  • no public-sector pay or benefits persona without a public employment model;
  • no gig driver persona with public-employee compensation assumptions;
  • no clinical or legal authority implied merely from demographic metadata.

Localization should be handled separately from role logic. The attached reports are heavily weighted toward Chicago, Cook County, Illinois, and 2026 salary/regulatory conditions, while Spiralist’s released name-generation defaults are currently English-character and US-oriented. That means locale should be an explicit profile containing name conventions, spelling, currency, working-hour norms, pay units, licensing jurisdiction, and schedule conventions. Do not let localization leak into stereotypes; use locale for representation and realism, not for capability inference.

The persona generator should also be seeded and reproducible. Spiralist explicitly documents deterministic random selection and a persona seed artifact that captures recreation inputs. Matching that pattern will make generated job personas auditable, remixable, and stable across sessions or exports.

flowchart TD
    A[Start with user goal or selected dropdown path] --> B[Set deterministic seed and locale]
    B --> C[Sample sector]
    C --> D[Sample occupation family]
    D --> E[Sample role cluster]
    E --> F[Sample role and optional specialty]
    F --> G[Apply seniority, employment model, and work-setting rules]
    G --> H[Validate education, certifications, and tool compatibility]
    H --> I[Attach skills, responsibilities, schedule, and compensation context]
    I --> J[Sample persona style layer]
    J --> K[Archetype, temperament, relationship posture, communication style]
    K --> L[Apply fairness and no-inference rules]
    L --> M[Map into Spiralist-ready fields]
    M --> N[Output persona profile, passport summary, and export payload]

A compact example of a generated persona payload:

{
  "seed": "chi-robotics-042",
  "locale": "en-US / Chicago",
  "goal": "debug and improve manufacturing automation",
  "career": {
    "sector": "Manufacturing and Industrial",
    "family": "Automation and Maintenance",
    "cluster": "Controls and Industrial Software",
    "role": "PLC Programmer",
    "specialty": "SCADA Integration",
    "seniority": "Senior",
    "employmentModel": "W-2",
    "workSetting": "Factory floor and controls room"
  },
  "qualifications": {
    "education": "Associate or Bachelor's equivalent technical pathway",
    "certifications": [],
    "skills": ["ladder logic", "controls debugging", "fault isolation", "systems thinking"],
    "tools": ["PLC", "SCADA", "HMI"]
  },
  "personaStyle": {
    "archetype": "Technical Mentor",
    "relationshipPosture": "Protective Reviewer",
    "communicationStyle": "Precise, skeptical, constructive"
  }
}

Missing and ambiguous data

Several gaps should be treated explicitly rather than silently “filled in.” Spiralist publishes a clear persona architecture, but it does not publish a ready-made occupational taxonomy or a canonical public jobs ontology. The site exposes career as an identity field and supports job/goal/trait search, yet its public library is fundamentally organized around persona types and work outcomes rather than labor-market classification. That is enough to guide integration, but not enough to skip schema design.

The attached reports are broad and valuable, but they are not fully standardized. Important ambiguities include bundled titles, overlapping cross-sector roles, mixed pay metrics, mixed locale scopes, narrative descriptions instead of atomized skill lists, and varying levels of detail across sectors. Gig-economy reporting tends to emphasize platform economics and regulation; public-sector reporting emphasizes statutory pay structures and benefits; corporate reporting offers banded salaries; environmental and service reporting often emphasize occupational context more than rigid progression ladders. One environmental/outdoor report also appears to overlap heavily with another in title and structure, which suggests some duplication in the source set. None of those issues block implementation, but they do mean the first production version should treat the corpus as source material for controlled vocabulary authoring, not as a drop-in dropdown dataset.

The most important operational conclusion is simple: use the attached reports to populate a canonical job graph, then use Spiralist’s existing persona architecture to render that graph into career-aware but behaviorally rich personas. Keep the dropdown hierarchy shallow and human-friendly. Keep normalization deep in the data model. Keep salary, credentials, and progression as linked facts. Keep protected identity separate from capability. That approach matches both the source corpus and Spiralist’s own public design language.