Α Β Γ Δ Ε

User Guide

How to use the Greek NLP Swiss Knife platform — ten research applications and approximately 200,000 lines of code for computational analysis of Greek texts from antiquity to the present.

← Back to Portal

Applications

Getting Started

The Greek NLP Swiss Knife is a suite of ten web applications for Greek text analysis, organized into four tiers based on their underlying technology. All applications run in your browser — no installation required. Some applications require API keys for the LLM providers they use.

API Keys: Applications in the LLM-powered and Neurosymbolic tiers require you to provide your own API keys from one or more of: Google Gemini, OpenAI, or Anthropic. These keys are entered directly in each application's interface and are never stored on the server. Classical NLP apps and the Dialect Generator do not require any API key.

Tier 1

Classical NLP

No LLM required — deterministic NLP pipelines for text analysis. No API key needed.

NATS

Narrative Analysis & Text System
Open App →

NATS provides named entity recognition with full Greek language support, document embeddings, interactive network visualizations, and community detection for textual analysis. It works entirely with classical NLP models, so no API key is required.

How to Use

  1. Open the app and paste or upload your Greek text.
  2. Choose the analysis type: Named Entity Recognition (NER), document embeddings, or network analysis.
  3. For NER, the system will identify and tag persons, locations, organizations, and other entities in your text.
  4. For network analysis, explore the interactive graph showing relationships between entities, including community detection clusters.
  5. Export results as needed for your research.

Named Entity Recognition

Identify persons, places, organizations in Greek text.

Network Visualization

Interactive graphs showing entity relationships.

Community Detection

Automatic clustering of related entities.

Document Embeddings

Vectorize texts for similarity comparison.

Voyant-NLP

Advanced Text Analysis & Embeddings Lab
Open App →

A comprehensive text analysis workbench inspired by Voyant Tools, enhanced with NLP capabilities specifically tuned for Greek. Upload one or more texts and explore them through multiple analytical lenses.

How to Use

  1. Upload one or more text files, or paste text directly into the input area.
  2. Browse the available analysis modules: word frequencies, concordance (KWIC), collocates, POS tagging, sentiment analysis, topic modeling, and more.
  3. For word embeddings, select Word2Vec or FastText training and explore word analogies and nearest neighbors.
  4. Use the NER module to extract named entities across your corpus.
  5. For multi-document analysis, use document clustering to find thematic groupings.

Word Frequencies

Frequency distributions with stopword filtering.

Concordance (KWIC)

Keyword in context search across your corpus.

Word2Vec / FastText

Train embeddings and explore word analogies.

Topic Modeling

Discover thematic clusters in your texts.

POS Tagging

Part-of-speech annotation for Greek.

Sentiment Analysis

Positive/negative polarity detection.

Tier 2

LLM-Powered

Multi-provider applications using Gemini, OpenAI, and Anthropic. Bring your own API key.

MEDEA-NEUMOUSA

Neuro-Symbolic AI for Classical Studies
Open App →

MEDEA is the most comprehensive application in the suite, offering six integrated AI capabilities for classical studies: cross-lingual translation across 18 ancient languages, knowledge graph extraction, emotional landscape analysis, and Prolog-based symbolic reasoning via the Zeugma module.

How to Use

  1. Enter your API key(s) — at least one of Google Gemini, OpenAI, or Anthropic.
  2. Select one of the six powers from the interface: Necromancer (translation), Knowledge Graphs, Emotional Landscape, Zeugma (logical reasoning), or others.
  3. Necromancer: Paste a text and select the source and target language pair. MEDEA supports 18 ancient and modern languages including Ancient Greek, Latin, Sanskrit, Hittite, and more.
  4. Knowledge Graphs: Paste a text and MEDEA will extract entities and relationships, visualizing them as an interactive knowledge graph.
  5. Emotional Landscape: Paste a literary text and MEDEA will analyze emotions using Cairns' theoretical framework, classifying them into 12 emotion families with Prolog-based symbolic verification.
  6. Zeugma: Input a logical argument in natural language and MEDEA will translate it into formal logic, verify it using Prolog, and report on its validity.
Tip: MEDEA supports multi-model comparison — you can run the same analysis across different LLM providers simultaneously to compare outputs.

Nature Analysis

Nature Discourse Detection in Text
Open App →

Analyze how nature is discussed in any text. The application segments text by paragraph, sentence, or token and uses LLMs to detect and classify nature-related discourse, producing statistical breakdowns and exportable data.

Login required: Use guest / nature123 to log in.

How to Use

  1. Log in with the guest credentials above, then enter your API key for at least one LLM provider.
  2. Paste or upload the text you want to analyze.
  3. Choose the segmentation level: paragraph, sentence, or token.
  4. Configure the detection threshold (how strictly nature discourse is identified).
  5. Run the analysis. The app will highlight nature-related segments and produce statistical summaries.
  6. Export results as CSV for further analysis in your own tools.
Δ

Linguistic Distance

Diachronic Evolution Measurement
Open App →

Measure linguistic distance across seven independent dimensions for eleven language pairs. This tool quantifies how languages have changed over time — from Ancient Greek to Modern Greek, Latin to Romance languages, and more — using established typological databases.

How to Use

  1. Select a language pair from the available options (e.g., Ancient Greek – Modern Greek, Latin – Italian).
  2. Choose which distance dimensions to compute: phonological (ASJP, Swadesh), morphosyntactic (UD Treebanks), typological (WALS, Grambank, URIEL+), and/or lexical.
  3. View the results as a multi-dimensional distance profile, showing how the two languages differ along each axis.
  4. Compare profiles across language pairs to understand relative rates of change.

Swadesh Lists

Core vocabulary comparison across time.

ASJP Distance

Automated phonological similarity.

UD Treebanks

Syntactic feature divergence.

WALS / Grambank

Typological feature comparison.

Tier 3

Neurosymbolic (NeSy)

LLM generation verified by Prolog reasoning, phonological engines, or rule-based parsers. API key required.

Greek Rhyme System

AI-Powered Rhyme Analysis & Generation
Open App →

Identify and generate rhyme patterns in Modern Greek poetry using a neurosymbolic pipeline that combines multi-model LLM analysis with a deterministic phonological engine. The system implements a complete taxonomy of Greek rhyme types and uses RAG-enhanced corpus retrieval for verification.

How to Use

  1. Enter your API key for at least one LLM provider.
  2. Rhyme Analysis: Paste a Greek poem or verse. The system will identify rhyme pairs, classify them by type (masculine/feminine, perfect/imperfect, etc.), and verify classifications using the phonological engine.
  3. Rhyme Generation: Provide a word or line ending, and the system will generate rhyming candidates verified against the Greek phonological ruleset.
  4. Explore the RAG corpus — the system retrieves similar rhyme patterns from a curated corpus of Greek poetry for comparison.
Neurosymbolic verification: The LLM proposes rhyme classifications, then a deterministic phonological analyzer verifies each classification against formal rules. If the LLM's classification conflicts with the phonological analysis, the system uses a Federative Loop to reconcile the disagreement.

PlotAnalyzer

Neuro-Symbolic Narrative Analysis
Open App →

Analyze the plot structure of any narrative text using five major narrative theories. PlotAnalyzer uses LLMs to identify structural elements and then verifies them against formal definitions of each theory's components.

How to Use

  1. Enter your API key for at least one LLM provider.
  2. Paste or upload a narrative text (novel excerpt, short story, dramatic script, etc.).
  3. Select one or more narrative theories to apply: Freytag's Pyramid (exposition, rising action, climax, falling action, denouement), Campbell's Hero's Journey, Booker's Seven Basic Plots, Aristotelian Poetics, or Russian Formalism (fabula vs. syuzhet).
  4. View the analysis: the system maps your text onto the chosen theory's structural elements and identifies key plot points, conflicts, and turning moments.
  5. Compare results across theories to see how different frameworks interpret the same narrative.

TermResearch AI

AI-Assisted Terminography with NeSy Validation
Open App →

Extract terms from domain-specific corpora, generate ISO 1087-compliant definitions using multi-LLM comparison, and validate each definition with a symbolic parser that checks for proper genus, differentia, circularity, and conciseness.

How to Use

  1. Enter your API key for at least one LLM provider.
  2. Upload a domain corpus (e.g., a set of texts from a specialized field like medicine, law, or technology).
  3. The system will extract candidate terms from your corpus using statistical and linguistic methods.
  4. For each term, TermResearch AI generates definitions using multiple LLMs, then validates each against ISO 1087 terminological standards using a symbolic parser.
  5. Review the validated definitions, and refine as needed. The parser flags issues like circular definitions, missing genus terms, or overly broad differentia.
ISO 1087 compliance: Every generated definition is checked by a rule-based parser to ensure it follows terminological best practices — proper genus-differentia structure, no circularity, adequate specificity, and appropriate conciseness.
Tier 4

Fine-Tuned Dialect Generation

Locally hosted models fine-tuned on the GRDD+ dialectal corpus. No API key needed.

Γ

Dialect Generator

Greek Dialectal Text Generation
Open App →

Generate text in four Greek regional dialects — Pontic, Cretan, Northern Greek, and Cypriot — using open-source language models fine-tuned with LoRA adapters on the GRDD+ dialectal corpus. Unlike the other applications, the Dialect Generator runs models locally on the server and requires no external API key.

How to Use

  1. Select a dialect from the four available options: Pontic (Ποντιακά), Cretan (Κρητικά), Northern Greek (Βόρεια Ελληνικά), or Cypriot (Κυπριακά).
  2. Select a model: Llama 3.1 8B Instruct, Llama 3 8B Instruct, or Krikri 8B Base. Each has been fine-tuned on the same dialectal data but may produce different stylistic results.
  3. Type your prompt in Greek — this is the text or topic you want the model to write about in the selected dialect.
  4. Adjust generation parameters if desired: Max tokens (10–256, default 80), Temperature (0.1–2.0, default 0.75), Top-p (0.1–1.0, default 0.90).
  5. Click Generate. Tokens will stream in one at a time. On CPU, generation takes roughly 5–10 seconds per token, so expect to wait 1–2 minutes for a typical output.
CPU inference: The models run on CPU (8 cores, 32 GB RAM) which makes generation slower than GPU. The first request after the app starts may take longer as the model loads into memory. Subsequent requests to the same model will be faster.

4 Dialects

Pontic, Cretan, Northern Greek, Cypriot.

3 Base Models

Llama 3.1, Llama 3, and Krikri 8B.

LoRA Fine-Tuned

Trained on 20k+ GRDD+ dialectal examples.

Streaming Output

Tokens appear as they are generated.

Frequently Asked Questions

Do I need to create an account?

No. All applications are freely accessible without registration. You only need API keys for the LLM-powered and Neurosymbolic tier applications.

Which LLM provider should I use?

All supported providers (Google Gemini, OpenAI, Anthropic) work well. For Greek-specific tasks, we recommend trying multiple providers using the multi-model comparison feature available in several applications. Each provider has different strengths with Greek text.

Are my texts stored on the server?

No. Texts are processed in memory during your session and are not stored. API keys are sent directly to the respective provider and are not retained.

Can I use these tools for Ancient Greek?

MEDEA-NEUMOUSA's Necromancer module supports Ancient Greek translation. NATS and Voyant-NLP can process any Greek text including Ancient Greek, though their NER models are primarily trained on Modern Greek. The Linguistic Distance tool includes Ancient Greek as a language pair.

The Dialect Generator is slow — is something wrong?

No. The dialect models run on CPU, which is inherently slower than GPU for neural text generation. Expect roughly 5–10 seconds per token. We are working on GPU acceleration for faster inference.

I get "Could not find a replica" or the app takes a long time to load.

Some applications — particularly the Dialect Generator — use scale-to-zero to manage hosting costs. When nobody has used the app for a while, Azure shuts down the container to save resources. The first request after that triggers a cold start: the container boots up and loads the model into memory, which can take a couple of minutes. Just wait and retry. Once the app is running, subsequent requests will be much faster (no restart needed).

Who built this platform?

The Greek NLP Swiss Knife was built by Stergios Chatzikyriakidis and collaborators. The dialect generation models are based on the GRDD+ dataset research. The full platform is described in the forthcoming DSH paper.