OPEN PROJECT
Multiscale analysis of texts – Developing a method for the complex characterization of a text’s structure based on multiscale analysis and neural-network theory
English version. Romanian version and Bibliography below.
Document retrieved from the archives of the UNESCO Chair of Geodynamics within the Institute of Geodynamics. The original text has been preserved as such, while the bibliography has been brought up to date through the addition of recent and relevant references.
Framework programme
The relationship between Creativity and the structure of language; the mother tongue as a process that particularizes a universal spiritual-phenomenological dimension.
Within this programme, the aims are:
– developing techniques and technologies for evaluating texts with a view to identifying semantic coherence and cultural/informative value;
– inter- and transdisciplinary studies on highlighting the “structure” of natural language (in comparison with artificial languages) and identifying cognitive and epistemic characteristics correlatable with the structure of language (with particular application to Romanian);
– comparative studies on the particularities of creativity and of the capacity for optimally approaching certain subjects as a function of the structure of a language; differences between natural languages and artificial languages.
1. Objectives
– developing an original methodology (the complex vectorization of the text, using for this purpose, among other things, fractal analysis) capable of individualizing a given text;
– the self-classification of texts through discrimination/clustering techniques and the study of possible correlations with their semantic content and with the psycho-emotional particularities of the author;
– objectifying the degree of **semantic coherence** of a work (an authentic original creation) by comparison with a translated text or one produced by collation;
– comparison between texts written in Romanian and in other languages but treating the same subject;
– providing ideas, concepts and techniques for formulating the themes and objectives of a study on evaluating the domains of performance in creativity/innovativeness as a function of the structure of a language (with applicability in forming high-performing multidisciplinary and multicultural teams).
2. Methods of realization
The analysis of texts through the multiscale method involves building a database containing a number of texts written in electronic format by several authors and applying an original methodology for their multiscale characterization, followed by the application of self-classification (clustering) techniques. The analysis will yield a self-grouping of the analysed texts according to vocabulary, syntactic and semantic particularities, allowing an objective classification.
The evaluation, interpretation and validation of the results obtained through the multiscale processing of the analysed texts will be supervised by a team of experts in the field of literary criticism, semantics and hermeneutics.
3. Expected results
The results expected to be obtained in this study will allow a better understanding of the subtle mental mechanisms involved in human creativity, of the role of language in this process, the objectification of expert assessments concerning the authenticity of texts, as well as a self-classification of texts according to their semantic content.
This is a preliminary stage for preparing the working methodology for approaching a major theme related to understanding the “role of the mother tongue in the mental process of creativity” (a process that allows one to grasp and develop new aspects of Reality), with applicability in enhancing human performance in the field of innovation.
PROIECT DESCHIS
Analiza multiscalară a textelor – Elaborarea unei metode de caracterizare complexă a structurii unui text pe baza analizei multiscalare și a teoriei rețelelor neurale
Document preluat din arhivele Catedrei de Geodinamică UNESCO din cadrul Institutului de Geodinamică. Textul original a fost păstrat ca atare, iar bibliografia a fost adusă la zi prin adăugarea unor referințe recente și relevante.
Program cadru
Relația dintre Creativitate și structura limbii; limba maternă ca proces ce particularizează o dimensiune spiritual-fenomenologică universală.
În cadrul acestui program se vizează:
– elaborarea de tehnici și tehnologii de evaluare a unor texte în vederea identificării coerenței semantice și a valorii cultural/informative;
– studii inter și transdisciplinare privind evidențierea „structurii” limbajului natural (comparație cu limbaje artificiale) și identificarea unor caracteristici cognitive și epistemice, corelabile cu structura limbii (particularizare pentru limba română);
– studii comparative privind particularități ale creativității și capacității de abordare optimală a unor subiecte în funcție de structura unei limbi; diferențe între limbaje naturale și limbaje artificiale.
1. Obiective
– elaborarea unei metodologii originale (vectorizarea complexă a textului utilizând în acest scop inclusiv analiza fractală) capabilă să individualizeze un text dat;
– auto-clasificarea textelor prin tehnici de discriminare/clusterizare și studiul unor posibile corelații cu conținutul lor semantic, respectiv cu particularități psiho-emoționale ale autorului;
– obiectivarea gradului de **coerență semantică** a unei lucrări (creație originală autentică) prin comparație cu un text tradus sau unul realizat prin colaționare;
– comparație între texte scrise în limba română și în alte limbi, dar care tratează același subiect;
– furnizarea de idei, concepte și tehnici pentru formularea tematicii și obiectivelor referitoare la un studiu privind evaluarea domeniilor de performanță în creativitate/inovativitate în funcție de structura unei limbi (cu aplicabilitate în formarea unor echipe multidisciplinare și multiculturale performante).
2. Modalități de realizare
Analiza textelor prin metoda multiscalară presupune realizarea unei bănci de date conținând un număr de texte scrise în format electronic de către mai mulți autori și aplicarea unei metodologii originale de caracterizare multiscalară a acestora, urmată de aplicarea unor tehnici de auto-clasificare (clusterizare). În urma analizei se va obține o auto-grupare a textelor supuse analizei în funcție de particularități de vocabular, sintactice și semantice, permițându-se o clasificare obiectivă.
Evaluarea, interpretarea și validarea rezultatelor obținute prin procesarea multiscalară a textelor analizate va fi supervizată de o echipă de experți în domeniul criticii literare, al semanticii și hermeneuticii.
3. Rezultate așteptate
Rezultatele preconizate a se obține în acest studiu vor permite o mai bună înțelegere a subtilelor mecanisme mentale implicate în creativitatea umană, a rolului limbajului în acest proces, obiectivarea expertizelor privind autenticitatea unor texte, precum și o auto-clasificare a textelor în funcție de conținutul semantic.
Aceasta este o etapă preliminară de pregătire a metodologiei de lucru pentru abordarea unei teme majore, legate de înțelegerea **rolului limbii materne în procesul mental al creativității** (proces ce permite a surprinde și dezvolta aspecte noi ale Realității), cu aplicabilitate în creșterea performanței umane în domeniul inovării.
Bibliografie / References
1. Shuklin, D. E. (2001). The structure of a semantic neural network extracting the meaning from a text. *Cybernetics and Systems Analysis*, 37(2).
2. Shuklin, D. E. (2001). The structure of a semantic neural network realizing morphological and syntactic analysis of a text. *Cybernetics and Systems Analysis*, 37(5).
3. Shuklin, D. E. (2002). Realization of a binary clocked linear tree and its use for processing texts in natural languages. *Cybernetics and Systems Analysis*, 38(4).
4. Dudar, Z. V., Shuklin, D. E. (2000). Implementation of neurons for semantic neural nets understanding texts in natural language. *Radio-electronika i Informatika*, KhTURE, 4, 89–96.
5. Shuklin, D. E. (2004). The further development of semantic neural network models. *Artificial Intelligence*, Institute of Artificial Intelligence „Nauka i obrazovanie”, Donețk, Ucraina, 3, 598–606.
6. Charniak, E. (2000). A maximum-entropy-inspired parser. În *Proceedings of the First Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)*, 132–139.
7. Gildea, D., Jurafsky, D. (2002). Automatic labeling of semantic roles. *Computational Linguistics*, 28(3), 245–288.
8. Manning, C. D., Schütze, H. (1999). *Foundations of Statistical Natural Language Processing*. MIT Press. ISBN 978-0262133609.
9. Zanette, D. H. (2006). Zipf’s law and the creation of musical context. *Musicae Scientiae*, 10(1), 3–18.
10. Kali, R. (2003). The city as a giant component: a random graph approach to Zipf’s law. *Applied Economics Letters*, 10, 717–720.
11. Gabaix, X. (1999). Zipf’s law for cities: an explanation. *Quarterly Journal of Economics*, 114(3), 739–767.
12. Privitera, C. M., Stark, L. W. (2000). Algorithms for defining visual regions-of-interest: comparison with eye fixations. *IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI)*, 22(9), 970–982.
13. Crosier, M. S., Griffin, L. D. (2007). Zipf’s law in image coding schemes. În *Proceedings of the British Machine Vision Conference (BMVC)*.
14. Sebeok, T. A. (1985). *Contributions to the Doctrine of Signs*. University Press of America, Lanham.
15. Pessa, E., Terenzi, G. (2007). Semiosis in cognitive systems: a neural approach to the problem of meaning. Springer, Berlin/Heidelberg.
16. Meystel, A. M. (2001). *Engineering of Mind: An Introduction to the Science of Intelligent Systems*. Wiley-Interscience.
17. Grossberg, S., Levine, D. S. (1987). Neural dynamics of attentionally modulated Pavlovian conditioning: blocking, inter-stimulus interval, and secondary reinforcement. *Psychobiology*, 15(3), 195–240.
18. Zadeh, L. A. (1997). Information granulation and its centrality in human and machine intelligence. În *Proceedings of the Conference on Intelligent Systems and Semiotics ’97*, Gaithersburg, MD, 26–30.
19. Grossberg, S., Schmajuk, N. A. (1987). Neural dynamics of attentionally modulated Pavlovian conditioning: conditioned reinforcement, inhibition, and opponent processing. *Psychobiology*, 15(3), 195–240.
20. Jurafsky, D., Martin, J. H. (2000). *Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition* (1st ed.). Prentice Hall.
21. Hosom, J.-P., Cole, R., Fanty, M., Schalkwyk, J., Yan, Y., Wei, W. (1999). *Training Neural Networks for Speech Recognition*. Center for Spoken Language Understanding, Oregon Graduate Institute of Science and Technology.
22. Lukanin, A. V. (2008). Linguistic synergetics and neural network approach (concerning R. G. Piotrovskii’s book).
23. Elliott, J., Atwell, E. (2000). Is anybody out there? The detection of intelligent and generic language-like features. *Journal of the British Interplanetary Society*, 53(1/2), 13–22.
24. Solé, R. V., Corominas-Murtra, B., Valverde, S., Steels, L. (2010). Language networks: their structure, function, and evolution. *Complexity*, 15(6), 20–26. Wiley Periodicals.
25. Roy, D. (2005). Semiotic schemas: a framework for grounding language in action and perception. *Artificial Intelligence*, 167(1–2), 170–205. Cognitive Machines Group, The Media Laboratory, MIT.
26. Christiansen, M. H., Kirby, S. (2003). Language evolution: consensus and controversies. *Trends in Cognitive Sciences*, 7(7), 300–307.
27. Chen, S., Goodman, J. (1996). An empirical study of smoothing techniques for language modeling. În *Proceedings of the 34th Annual Meeting of the Association for Computational Linguistics (ACL)*.
28. Deely, J. (2003). *The Impact on Philosophy of Semiotics*. St. Augustine’s Press, South Bend.
29. Deely, J. (2001). *Four Ages of Understanding*. University of Toronto Press, Toronto.
30. Steels, L. (2007). Language as a complex adaptive system. Bruxelles.
31. Steels, L. (1997). The synthetic modeling of language origins. *Evolution of Communication*, 1(1), 1–34.
32. Kaplan, F. (1999). Dynamiques de l’auto-organisation lexicale: simulations multi-agents et „Têtes parlantes”. *In Cognito*, 15, 3–23.
33. McIntyre, A. (1998). Babel: a testbed for research in the origins of language. În *Proceedings of COLING-ACL ’98*, Montréal.
34. Drożdż, S., Oświęcimka, P., Kulig, A., Kwapień, J., Bazarnik, K., Grabska-Gradzińska, I., Rybicki, J., & Stanuszek, M. (2016). Quantifying origin and character of long-range correlations in narrative texts. *Information Sciences*, 331, 32–44. https://doi.org/10.1016/j.ins.2015.10.023
35. Stanisz, T., Drożdż, S., & Kwapień, J. (2024). Complex systems approach to natural language. *Physics Reports*, 1053, 1–84. https://doi.org/10.1016/j.physrep.2023.12.002
36. Chatzigeorgiou, M., Constantoudis, V., Diakonos, F., Karamanos, K., Papadimitriou, C., Kalimeri, M., & Papageorgiou, H. (2017). Multifractal correlations in natural language written texts: Effects of language family and long word statistics. *Physica A: Statistical Mechanics and its Applications*, 469, 173–182. https://doi.org/10.1016/j.physa.2016.11.028
37. Cong, J., & Liu, H. (2014). Approaching human language with complex networks. *Physics of Life Reviews*, 11(4), 598–618. https://doi.org/10.1016/j.plrev.2014.04.004
38. Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. *Proceedings of the International Conference on Learning Representations (ICLR), Workshop Track*. arXiv:1301.3781. https://arxiv.org/abs/1301.3781
39. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. *Advances in Neural Information Processing Systems (NeurIPS)*, 30, 5998–6008. arXiv:1706.03762. https://arxiv.org/abs/1706.03762
40. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. *Proceedings of NAACL-HLT 2019*, 4171–4186. https://doi.org/10.18653/v1/N19-1423
41. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. *Proceedings of EMNLP-IJCNLP 2019*, 3982–3992. https://doi.org/10.18653/v1/D19-1410
42. Huang, B., Chen, C., & Shu, K. (2024). Authorship attribution in the era of LLMs: Problems, methodologies, and challenges. *ACM SIGKDD Explorations Newsletter*, 26(2), 21–43. https://doi.org/10.1145/3715073.3715076
43. Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero-shot machine-generated text detection using probability curvature. *Proceedings of the 40th International Conference on Machine Learning (ICML)*, PMLR 202, 24950–24962. arXiv:2301.11305. https://arxiv.org/abs/2301.11305
