Posts tonen met het label corpus. Alle posts tonen
Posts tonen met het label corpus. Alle posts tonen

dinsdag 27 mei 2025

corpus social media NL

 

Ja, er bestaan verschillende corpora en bronnen waar je Nederlandse gesprekken kunt vinden, ook in de persoonlijke sfeer. Hier zijn enkele relevante opties:


🇳🇱 Nederlandstalige corpora met gesprekken

1. Corpus Gesproken Nederlands (CGN)

  • Beschrijving: Groot corpus van gesproken Nederlands uit Nederland en Vlaanderen.

  • Inhoud: Bevat o.a. informele gesprekken, telefoongesprekken, interviews, discussies en meer.

  • Persoonlijke sfeer?: Ja, vooral in de informele dialogen (bijv. spontane gesprekken tussen vrienden of familie).

  • Toegang: Vereist registratie. Zie: https://ivdnt.org/taalmaterialen/cgn/

2. CHILDES - Dutch corpora

  • Beschrijving: Taalontwikkelingscorpora, inclusief ouder-kind-interacties in het Nederlands.

  • Inhoud: Persoonlijke gesprekken tussen ouders/verzorgers en kinderen.

  • Toegang: Vrij toegankelijk via: https://childes.talkbank.org/ → Zoek op Dutch.

3. TALP corpora (Tilburg University)

  • Bevat conversatiecorpora met specifieke settings (bijv. gesprekken tussen zorgverleners en patiënten).

  • Zie: https://cls.ru.nl/ en zoek naar “gesproken taalcorpora”.

4. OpenSoNaR

  • Beschrijving: Groot corpus geschreven en gesproken Nederlands, inclusief blogs en social media (meer informele taal).

  • Minder gesproken taal, maar informele sfeer.

  • https://opensonar.ivdnt.org/


📚 Andere suggesties

5. YouTube + automatische transcriptie

Gebruik Nederlandstalige vlogs, interviews of podcasts en genereer automatische transcripties via YouTube of een tool zoals autosub of Whisper van OpenAI.

6. Reddit en fora (zoals Viva Forum of Ouders.nl)

  • Persoonlijke gesprekken (al zijn het geschreven dialogen).

  • Je kunt scraping tools gebruiken om deze te verzamelen, maar let op ethische en privacyrichtlijnen.


Zoek je dit voor onderzoek, taalanalyse of iets anders?

Als je meer context geeft (bijv. je wil zelf een dataset bouwen, je zoekt voorbeelden voor onderwijs, etc.), kan ik gerichtere bronnen of hulp bieden.

woensdag 18 januari 2012

automatische kwaliteitsbepaling van forumposts

Assessing the quality of user generated content is an important problem for many web forums. While quality is currently assessed manually, we propose an algorithm to assess the quality of forum posts automatically and test it on data provided by Nabble.com. We use state-of-the-art classification techniques and experiment with five feature classes: Surface, Lexical, Syntactic, Forum specific and Similarity features. We achieve an accuracy of 89% on the task of automatically assessing post quality in the software domain using forum specific features. Without forum specific features, we achieve an accuracy of 82%.


Clues for detecting irony in user-generated contents: oh...!! it's "so easy" ;-)

We investigate the accuracy of a set of surface patterns in identifying ironic sentences in comments submitted by users to an on-line newspaper. The initial focus is on identifying irony in sentences containing positive predicates since these sentences are more exposed to irony, making their true polarity harder to recognize. We show that it is possible to find ironic sentences with relatively high precision (from 45% to 85%) by exploring certain oral or gestural clues in user comments, such as emoticons, onomatopoeic expressions for laughter, heavy punctuation marks, quotation marks and positive interjections. We also demonstrate that clues based on deeper linguistic information are relatively inefficient in capturing irony in user-generated content, which points to the need for exploring additional types of oral clues.

Topic-sentiment analysis for mass opinion

corpusonderzoek en automatisering

Automatic creation of a reference corpus for political opinion mining in user-generated content



We propose and evaluate a method for automatically creating a reference corpus for training text classification procedures for mining political opinions in user-generated content. The process starts by compiling a collection of highly opinionated comments posted by users on an on-line newspaper. Then, we define and use a set of manually-crafted high-precision rules supported by a large sentiment-lexicon in order to identify sentences in each comment expressing opinions about political entities. Finally, the opinions found are propagated to the remainder sentences of the comment mentioning the same entities, thus increasing the number and variety of opinion-bearing sentences. Results show that most of the rules can identify negative opinions with very high precision, and these can be safely propagated to the remainder sentences in the comment in almost 100% of the cases. Due to problems arising from irony, the precision of identification drops for positive opinions, but several rules still reach high precision. Propagation of positive opinions is correct in about 77% of the cases, and most errors at this stage result from irony and polarity inversion throughout the comment.