Migration to Semantic Device Identifiers
| Cite this page | |
|---|---|
| 🔗 DOI | 10.64393/balagha.semantic.ID |
| Date | 10 February 2026 |
| Supplementary file | Download from Github |
Background
When the Encyclopedia of Arabic Rhetoric was launched in 2022, rhetorical devices were assigned alphanumeric identifiers such as A-5, B-2, and CH-1. In this system:
- The first alphabetic component referred to the domain of classical Arabic rhetoric the device was located within: one of ‘ilm al-ma‘ānī (Domain A), ‘ilm al-bayān (Domain B), and ‘ilm al-badī‘ (Domain C).
- For devices in ‘ilm al-badī‘ (Domain C), a second alphabetic component referred to one of 10 thematic subcategories created by the Encyclopedia for ease of reference.
- The numeric components were allocated sequentially.
These identifiers were introduced primarily for practical reasons: they made it easier to locate and reference individual devices within an initial collection of 95 rhetorical devices. At this stage, the alphanumeric identifier system functioned as a convenient indexing mechanism rather than as a theoretically-motivated classification scheme.
Limitations of the numeric system
Over time, several limitations of the original alphanumeric identifiers became apparent.
1) Implied Hierarchy
The use of sequential and domain-based numbering unintentionally suggested a hierarchical or developmental ordering between rhetorical devices. In practice, however, no such hierarchy exists within Arabic rhetorical theory. Devices are related in complex and overlapping ways, and their relative importance cannot be reduced to a numerical position.
2) Structural rigidity
Because identifiers were tied to specific domains and positions, the system restricted the reorganisation of devices. Moving a device from one domain to another, or from one subcategory to another, would disrupt the numbering sequence and create inconsistencies. As new subcategories were developed, it was impossible to relocate rhetorical devices appropriately due to the rigidity of the alphanumeric identifier system. As a result, the taxonomy became increasingly resistant to revision, despite ongoing scholarly refinement.
3) Limited memorability of numeric identifiers
The alphanumeric identifier system was a hindrance to pedagogy and operationalisation of the taxonomy because each identifier was opaque: a user cannot infer the meaning of an identifier like CE-3.
4) Limited cross-domain portability
An identifier like CH-11 is implicitly restricted to the Encyclopedia. It references rhetorical device number 11, within subcategory H, in domain C: a hierarchy which is unique to this Encyclopedia, but which is meaningless when ported to a different application such as the BALAGHA Score, or the Balagha Corpus.
The February 2026 Taxonomy Revision
In February 2026, a major phase of reorganisation was undertaken. This included:
- The refinement and subdivision of each domain - A, B, and C.
- The relocation of certain devices to more appropriate theoretical neighbourhoods.
- The clarification of relationships between overlapping categories.
These changes made it clear that the existing alphanumeric identifier system was no longer suitable for a developing and flexible taxonomy.
Introduction of semantic identifiers
To support long-term portability and structural flexibility, the alphanumeric identifiers were replaced with semantic identifiers in February 2026.
Each device is now assigned a stable alphabetic identifier of three to seven letters (for example, CONJUN: Conjunction, DEFINT: Definiteness, and VERBOS: Verbosity). These identifiers are derived from the English name of the device and are intended to be mnemonic and recognisable.
While the term "semantic" is used here in a broad sense, it reflects the fact that the new identifiers are meaning-based rather than position-based. They encode conceptual identity rather than taxonomic location.
Advantages of the new system
1) Conceptual neutrality
Semantic identifiers do not imply hierarchy, sequence, or priority. They function purely as labels, allowing devices to be treated as theoretically independent units.
2) Structural portability
Because identifiers are no longer tied to domains or subcategories, devices can be moved, merged, or subdivided in the future without requiring changes to their core references. This supports ongoing refinement of the taxonomy.
3) Improved readability and recall
Alphabetic identifiers based on device names are easier to remember and interpret than abstract alphanumeric sequences. This benefits researchers, students, and tool developers working with the Encyclopedia.
4) Technical compatibility with external applications
Stable, meaning-based identifiers are better suited to computational processing, corpus annotation, and cross-platform integration, particularly within the wider BALAGHA Score ecosystem.
Transitional policy
During the migration process, legacy alphanumeric identifiers have been documented in page histories and redirects have been maintained where appropriate. This ensures continuity for existing citations and external references. Where a former composite device has been split into multiple independent entries, this relationship is recorded in the relevant page development sections.
Ongoing development
The migration to semantic identifiers forms part of a broader effort to establish durable editorial standards for the Encyclopedia of Arabic Rhetoric. The system may be refined in the future, but any substantial changes will be documented in accordance with established versioning and editorial policies. The primary objective remains to balance theoretical accuracy, technical robustness, and long-term sustainability.
Mapping between old and new identifiers
This file on Github provides a full mapping between the old and new identifiers.
Page history
- 2026-02-10 - Added DOI and link to mapping file.
- 2026-02-09 - Page created: Migration to Semantic Device Identifiers.