A hands-on guide to Semantic Tagger for your text data analysis

22 March 2023

The Australian Text Analytics Platform (ATAP) project is a project that aims to provide researchers with the tools and training for analysing, processing and exploring text. As part of this project, we have adapted, with permission, a Semantic Tagger developed by the University Centre for Computer Corpus Research on Language (UCREL) at Lancaster University. This tool uses the Python Multilingual UCREL Semantic Analysis System (PyMUSAS) to tag your text data so that you can extract token level semantic tags from your text.

In addition to the USAS tags, this tool can also recognise Multi Word Expressions (MWE), i.e. expressions formed by two or more words that behave like a unit such as 'South Australia', and identify lemmas and Part-of-Speech (POS) tags in the text. For example, in the sentence ‘President Joe Biden attended two meetings today’, the tool will tag each token with its semantic tag like this -> ‘President Joe Biden’: MWE of [Personal names], ‘attended’: [Participating], ‘two’: [Number], ‘meetings’: [Participating] and ‘today’: [Time: Present; simultaneous]. This tool is available in both English and multi-lingual (Chinese, Italian and Spanish) versions and supports saving the results locally for further analysis, enabling you to gain meaningful insights into your research questions.

Length: 90 minutes

Leaders: Sony Jufri