Student projects

This page lists the projects currently open at the Laboratory for the History of Science and Technology (LHST). If you are interested in working on one of the projects listed below, please write directly to the contact person, with Prof. Baudry cc’ed in your email.

The project descriptions are only brief outlines and we are in general flexible about the particulars.

This project, in collaboration with the Institut des humanités en médecine (IHM), is part of the SNSF-funded research project MEDIF, led by Dr. Aude Fauvel and Prof. Rémy Amouroux. The aim of MEDIF is to examine the collective contribution of the first French and French-speaking Swiss women doctors to the development of the theoretical and practical frameworks of medicine between 1870 and 1940.

The proposed project examines the use of illustrations included by doctors in the health manuals they have written. The aim is to understand the body being treated from the perspective of both female and male doctors.

The corpus consists of about eighty health manuals written in French and published in both Switzerland and France between 1870 and 1940. These works are available in digitized form on library websites, with varying levels of OCR quality. The aim of the project is to extract the images, identify the type of image (photograph, drawing, etc.), what they depict (machine, part of the body, child, adult), and the number per volume. We will then focus on images depicting human beings, to examine the role of each gender in medical discourse.

This project will be carried out in collaboration with Dre. Amélie Puche (IHM) and Prof. Jérôme Baudry (EPFL).

Project type: semester project or master’s thesis.

Prerequisites: Prior experience in computer vision; solid skills in data analysis and Python;  interest in history and social sciences a plus; language skills in French.

The project investigates how digital tools can be used to study the development of botany, forestry, and agricultural knowledge in 19th-century France through the case of the acclimatization of Eucalyptus, an Australian plant, discovered in 1788 by Charles Louis L’Héritier. 

The corpus includes 3,660 books published between 1791 and 1914 that mention eucalyptus, available in bulk downloads with good OCR quality. These are books on medicine, agriculture, botany, history, geography, and literature available on Gallica, the online platform of the French national library (BnF). 

The project aims to determine, with the use of text mining, the evolution of publications dealing with the eucalyptus through the 19th century and identify the type of properties this tree is associated with (e.g. medicinal, forestry, agricultural, cleansing). To do this, the student will have to 

  • harvest the corpus through Gallica’s API;
  • assemble it into a comprehensive DataFrame with relevant metadata;
  • create a database and carry out textual analysis, which can range:

    a) co-occurrence networks;
    b) topic-modeling;
    c) sentiment analysis;
    d) diachronic evolution of issues and topics associated with eucalyptus;
    e) possibly, named entity recognition to construct knowledge graphs. 

This student project is part of a broader environmental history research project: the study of environmental, technical, scientific, and social transformations brought in by the planting of Eucalyptus, a “new” tree in the second half of the nineteenth and twentieth centuries in the Mediterranean region.

Project type: semester project or master’s thesis.

Prerequisites: Prior experience in text mining; solid data analysis skills; knowledge in LLM a plus; interest in environmental history a plus; language skills in French.

Possibility to work in group? Yes.

Contact: Elisabeth Davin-Mortier ([email protected])

In brief: This project aims at extracting data from digitized historical trade statistics in order to analyze the environmental footprint of Switzerland over the 20th century.

Context: In current discussions on environmental issues, materials have received increasing attention. The unequal distribution of the consequences of their extraction have been highlighted and criticized, for instance in the case of lithium for batteries. Similarly, scholars and activists have pointed to the misleading impression given by statistics that show a decline of CO2 emissions in some high-income countries in recent decades: while less carbon is emitted inside one country, more is emitted to produce and transport the imported goods.

There are well-established methods to investigate such questions, including material flow analysis, that considers all the flows and stocks of materials within a space, e. g. a city, a region or a country. It relies on data on local extraction (including mines and quarries, fishery, wood and crops), on imports and exports, on waste and on emissions. The method most often focuses on current material flows. National analyses usually start only in the 1970s, and long term change has been reconstructed only for a dozen countries.

Objective: The aim of the project is to start contributing to such a longer-term analysis for Switzerland, using digitized trade statistics from 1890 to 2000. These official statistics report the quantity and monetary value of the imports and exports of a given commodity, broken down by country of origin or destination.

The student will first build a pipeline to extract data from the scans. This includes determining an evaluation method, establishing a baseline and running experiments to guarantee good results. Next, the extracted data has to be integrated into a coherent database. Some of the most interesting challenges in these steps reside in the table format of the statistics and the changing categorization of the commodities.

Ideally, the project would then examine one research question in light of this data, depending on the student’s interests. For example, the project could try to assess how the carbon footprint of the country changed over time; examine the evolution of its imports of food and fodder; or investigate some critical materials, such as coal and oil.

Project type: semester project or master’s thesis.

Prerequisites: Solid programming and data analysis skills, and familiarity with machine learning workflows are required. Experience with LLMs and computer vision methods are desirable. Experience with historical corpora, OCR methods, or digital humanities methods would be an advantage.

 

Open Science is an international movement aiming at making all scientific research productions—publications, data, software, methods—freely accessible to all people in society: researchers, amateurs, policy makers, industries, as well as artists, journalists, and activists. Open Science across the world relies heavily on the design and development of dedicated infrastructures, mostly platforms: digital libraries, data repositories (“as open as possible, as closed as necessary”), directories, online journals, web archives, computational services, MOOCs, content management systems (CMS), collaborative version control, etc.

These platforms are loosely bound as a network. For example, a publication on an online journal may refer, via a persistent identifier (such as a DOI), to a dataset hosted on a given repository. Another example is when an open science search engine may harvest the metadata of libraries and directories to index available publications.

The shape of the open science ecosystem online and the nature of the links that tie platforms together are not well known yet. The aim of this project is to crawl the web to identify the links between platforms, to characterize their nature, and to generate an interactive map of the open science network.

Level
Master (research project, optional research project, or master’s project).

Possibility to work in group?
Yes.

Contact: Simon Dumas Primbault ([email protected])

For some years now, the term ‘ecosystem’ has been used by a number of research stakeholders to refer, in very different ways, to the environment in which their practice takes place. The numerous uses of the semantic field of ecology to understand the digital transition of research environments are, however, very diverse and very polarised.

The first part of this projet will be to assemble a vast heterogeneous corpus comprising writings of very different genres (scientific articles, policy briefs, reports, tribunes) as well as oral documents (courses, conferences, speeches) and metadata alone. This corpus will be built up by harvesting targeted sources: databases of scientific articles, crawling of research networks and infrastructures, archives of administrative institutions, course and conference repositories. Note: for the first instantiation of the project, a smaller and less heterogeneous corpus can be assembled.

From there, the project can take two different but complementary directions:

  1. Use NLP to perform an automated systematic literature review in order to trace the emergence of the term, and other relevant semantic fields, their circulation and their crystallization.
  2. Use network analysis in order to create a multi-layered network of documents, authors, concepts, institutions and disciplines and produce a diachronic map of the emergence and circulation of ecological discourse on research.

Project type: Master (research project, optional research project, or master’s project).

Possibility to work in group? Yes.

Contact: Simon Dumas Primbault ([email protected])