Exploring the rise of green growth issues within the EU and the expertise underlying the European Green New Deal
R 100%
<1%

README.md

EU and Green Growth #

Aurélien Goutsmedt

Project overview #

This repository contains the data workflow and text analysis for a research paper on the expertise underlying the European Green Deal, with a focus on the “green growth” concept.

The empirical strategy compares two institutional spaces:

  • DG ECFIN (European Commission) publications
  • JRC (Joint Research Centre) publications and metadata

The project combines document preprocessing, exploratory analysis, and Structural Topic Modelling (STM) to map themes, their prevalence, and their evolution over time.

Research focus #

The core objective is to identify how “green growth” is framed, circulated, and transformed across:

  1. policy-oriented economic documents (ECOFIN), and
  2. research-oriented technical production (JRC).

This comparison helps characterize differences in vocabulary, thematic emphasis, and temporal dynamics of expertise.

Repository structure #

.
├── eu_green_growth.Rproj
├── packages_and_data_path.R
├── documents/
│   └── exploring_documents.qmd
├── figures/
└── R/
    ├── clean_jrc_data.R
    ├── analyse_jrc_data.R
    ├── prepare_topic_model.R
    ├── run_topic_model.R
    ├── analysing_topic_model.R
    ├── background_topic_model.R
    ├── helper_functions.R
    └── depreciated/
        └── preparing_topic_model.R

Main scripts and roles #

  • packages_and_data_path.R
    Loads packages, defines local data paths, and sources helper functions.

  • R/clean_jrc_data.R
    Builds cleaned JRC metadata and identifies economics-related documents.

  • R/analyse_jrc_data.R
    Explores JRC corpus composition (collections, science areas, keyword trends, “green growth” signals).

  • R/prepare_topic_model.R
    Prepares combined ECOFIN + JRC text corpus, extracts n-grams, computes TF-IDF, and saves filtered tokens.

  • R/run_topic_model.R
    Builds STM input objects, fits STM models, and estimates topic prevalence effects over time.

  • R/analysing_topic_model.R
    Produces topic diagnostics and figures (top terms, prevalence, time evolution, topic correlation network).

  • documents/exploring_documents.qmd
    Narrative exploration document with descriptive outputs and selected visualizations.

  • R/helper_functions.R
    Utility functions for feature filtering, STM diagnostics, topic labelling, and visualization helpers.

Data #

The project uses local data directories configured in packages_and_data_path.R:

  • ECOFIN source data
  • JRC source data
  • project-level derived data

Because these paths are local and data files are not fully versioned in this repository, reproducibility requires access to the same (or equivalent) data sources.

Typical workflow #

  1. Configure package loading and paths in packages_and_data_path.R.
  2. Clean and filter JRC metadata: R/clean_jrc_data.R.
  3. Prepare tokens and corpus: R/prepare_topic_model.R.
  4. Fit STM and save model outputs: R/run_topic_model.R.
  5. Analyze topics and generate figures: R/analysing_topic_model.R.
  6. Explore and present results in Quarto: documents/exploring_documents.qmd.

Outputs #

Main outputs include:

  • cleaned metadata (*.rds)
  • token tables for topic modelling
  • fitted STM objects and effect estimations
  • figures in figures/ (topic prevalence, top terms, temporal effects)

Notes #

  • The repository currently mixes active and older scripts (including R/depreciated/).
  • Some script comments still reference legacy paths/filenames; current workflow should follow the script list above.
  • The project is research-in-progress and can be extended with robustness checks (alternative topic counts, alternative token filtering thresholds, and cross-corpus comparison diagnostics).