Migration to Semantic Device Identifiers

    From The Encyclopedia of Arabic Rhetoric
    Cite this page
    🔗 DOI10.64393/balagha.semantic.ID
    Date10 February 2026
    Supplementary fileDownload from Github

    Background

    When the Encyclopedia of Arabic Rhetoric was launched in 2022, rhetorical devices were assigned alphanumeric identifiers such as A-5, B-2, and CH-1. In this system:

    These identifiers were introduced primarily for practical reasons: they made it easier to locate and reference individual devices within an initial collection of 95 rhetorical devices. At this stage, the alphanumeric identifier system functioned as a convenient indexing mechanism rather than as a theoretically-motivated classification scheme.

    Limitations of the numeric system

    Over time, several limitations of the original alphanumeric identifiers became apparent.

    1) Implied Hierarchy

    The use of sequential and domain-based numbering unintentionally suggested a hierarchical or developmental ordering between rhetorical devices. In practice, however, no such hierarchy exists within Arabic rhetorical theory. Devices are related in complex and overlapping ways, and their relative importance cannot be reduced to a numerical position.

    2) Structural rigidity

    Because identifiers were tied to specific domains and positions, the system restricted the reorganisation of devices. Moving a device from one domain to another, or from one subcategory to another, would disrupt the numbering sequence and create inconsistencies. As new subcategories were developed, it was impossible to relocate rhetorical devices appropriately due to the rigidity of the alphanumeric identifier system. As a result, the taxonomy became increasingly resistant to revision, despite ongoing scholarly refinement.

    3) Limited memorability of numeric identifiers

    The alphanumeric identifier system was a hindrance to pedagogy and operationalisation of the taxonomy because each identifier was opaque: a user cannot infer the meaning of an identifier like CE-3.

    4) Limited cross-domain portability

    An identifier like CH-11 is implicitly restricted to the Encyclopedia. It references rhetorical device number 11, within subcategory H, in domain C: a hierarchy which is unique to this Encyclopedia, but which is meaningless when ported to a different application such as the BALAGHA Score, or the Balagha Corpus.

    The February 2026 Taxonomy Revision

    In February 2026, a major phase of reorganisation was undertaken. This included:

    • The refinement and subdivision of each domain - A, B, and C.
    • The relocation of certain devices to more appropriate theoretical neighbourhoods.
    • The clarification of relationships between overlapping categories.

    These changes made it clear that the existing alphanumeric identifier system was no longer suitable for a developing and flexible taxonomy.

    Introduction of semantic identifiers

    To support long-term portability and structural flexibility, the alphanumeric identifiers were replaced with semantic identifiers in February 2026.

    Each device is now assigned a stable alphabetic identifier of three to seven letters (for example, CONJUN: Conjunction, DEFINT: Definiteness, and VERBOS: Verbosity). These identifiers are derived from the English name of the device and are intended to be mnemonic and recognisable.

    While the term "semantic" is used here in a broad sense, it reflects the fact that the new identifiers are meaning-based rather than position-based. They encode conceptual identity rather than taxonomic location.

    Advantages of the new system

    1) Conceptual neutrality

    Semantic identifiers do not imply hierarchy, sequence, or priority. They function purely as labels, allowing devices to be treated as theoretically independent units.

    2) Structural portability

    Because identifiers are no longer tied to domains or subcategories, devices can be moved, merged, or subdivided in the future without requiring changes to their core references. This supports ongoing refinement of the taxonomy.

    3) Improved readability and recall

    Alphabetic identifiers based on device names are easier to remember and interpret than abstract alphanumeric sequences. This benefits researchers, students, and tool developers working with the Encyclopedia.

    4) Technical compatibility with external applications

    Stable, meaning-based identifiers are better suited to computational processing, corpus annotation, and cross-platform integration, particularly within the wider BALAGHA Score ecosystem.

    Transitional policy

    During the migration process, legacy alphanumeric identifiers have been documented in page histories and redirects have been maintained where appropriate. This ensures continuity for existing citations and external references. Where a former composite device has been split into multiple independent entries, this relationship is recorded in the relevant page development sections.

    Ongoing development

    The migration to semantic identifiers forms part of a broader effort to establish durable editorial standards for the Encyclopedia of Arabic Rhetoric. The system may be refined in the future, but any substantial changes will be documented in accordance with established versioning and editorial policies. The primary objective remains to balance theoretical accuracy, technical robustness, and long-term sustainability.

    Mapping between old and new identifiers

    This file on Github provides a full mapping between the old and new identifiers.

    Page history