The Role of Global University Benchmarking in Advancing Water Sustainability and Management Research
- 19 hours ago
- 26 min read
Author: Sofia García
Affiliation: Swiss International University (SIU)
ORCID ID: 0009-0009-5070-1669
Submitted 21 February 2026; Revised 09 May 2026; Revised 22 June 2026; Accepted 02 July 2026; Available online 08 August 2026; Version of Record 08 August 2026.
Doi: https://doi.org/10.65326/u7y.SpecSDG10005
Volume 3, December 2026, (SpecSDG10005)

Abstract
Global benchmarking instruments now shape how universities present, organise, and fund their sustainability work, yet their consequences for water-focused research remain poorly understood. This article develops an integrative review of three literatures that have grown largely in isolation: the sociology and scientometrics of university rankings, sustainability assessment in higher education, and research on water sustainability and Sustainable Development Goal 6. Drawing on peer-reviewed studies published mainly between 2003 and 2022, together with the published methodologies of the Times Higher Education Impact Rankings, UI GreenMetric, and the Sustainability Tracking, Assessment and Rating System, the review examines how water is represented in benchmarking architectures, through which institutional mechanisms participation can plausibly strengthen water research capacity, and where benchmarking threatens to distort that capacity. The analysis shows that water occupies a structurally weak position in most instruments: it is absent from academic world rankings, optional in the Impact Rankings, and folded into campus operations elsewhere. A conceptual framework is proposed that links benchmarking participation to water research capacity through three mechanisms, namely signalling and resource mobilisation, measurement and data infrastructure, and network formation, each conditioned by indicator validity and by reactive institutional behaviour. The framework yields testable propositions and design principles for ranking stewards, university leaders, and funders who want benchmarking to serve, rather than substitute for, substantive water scholarship.
Keywords: university rankings, water sustainability, sustainable development goals, higher education, research assessment, benchmarking, SDG 6
1. Introduction
Freshwater systems are under documented and intensifying pressure. Vörösmarty et al. (2010) mapped concurrent threats to human water security and river biodiversity at the global scale, and Mekonnen and Hoekstra (2016) estimated that four billion people face severe water scarcity. Hoekstra and Mekonnen (2012) traced the water footprint of humanity across production and consumption chains, and Gleick (2003) argued that twenty-first-century water challenges demand demand-side, efficiency-oriented solutions rather than supply expansion alone. These conditions define a research agenda that universities are uniquely positioned to advance: they train water professionals, operate campuses with measurable water demands, and produce the hydrological, engineering, and governance scholarship on which Sustainable Development Goal 6 (SDG 6, clean water and sanitation) depends.
At the same time, universities increasingly organise their sustainability activity around external benchmarking instruments. The Times Higher Education (THE) Impact Rankings score institutions against the seventeen SDGs, including a dedicated SDG 6 pillar (Times Higher Education, 2024b). UI GreenMetric ranks campuses on criteria that include water usage (UI GreenMetric, n.d.). The Sustainability Tracking, Assessment and Rating System (STARS) awards credits for water performance within its operations category (Association for the Advancement of Sustainability in Higher Education [AASHE], n.d.). Participation in such systems has become a strategic act: it signals commitment to funders and applicants, structures internal reporting, and, as a long line of scholarship on rankings shows, changes the behaviour of the organisations being measured (Espeland & Sauder, 2007).
The difficulty is that these two developments are studied by communities that rarely meet. Scholars of rankings have produced a rich critique of indicator validity and organisational reactivity (Gadd, 2021; Marginson, 2014; Selten et al., 2020), but they seldom examine environmental content, and almost never water specifically. Scholars of sustainability in higher education have compared assessment tools and documented implementation gaps (Alghamdi et al., 2017; Galleli et al., 2022; Lozano et al., 2015), but they treat water as one operational category among many rather than as a research field with its own capacity requirements. Water scholars, for their part, have mapped SDG-related publication trends (Salvia et al., 2019; Sianes et al., 2022) without asking how benchmarking participation feeds back into the research agendas they measure. The result is a precise and consequential gap: there is no integrated account of how global university benchmarking represents water sustainability, through which mechanisms participation could strengthen or weaken water research capacity, and what design features would make the difference. The gap matters for ranking stewards deciding how to weight and verify water indicators, for university leaders deciding whether SDG 6 participation is worth its reporting cost, and for funders and water agencies deciding whether ranking positions carry any signal about research capability.
This article addresses that gap through an integrative review with a conceptual contribution. Following Torraco (2016) and Whittemore and Knafl (2005), it synthesises heterogeneous literatures in order to generate a new framework rather than to aggregate effect sizes. Three research questions structure the analysis:
· RQ1. How do global university benchmarking systems represent water sustainability and management in their indicator architectures, and with what validity?
· RQ2. Through which institutional mechanisms can benchmarking participation plausibly shape universities' water-related research capacity and agendas?
· RQ3. Which unintended consequences of benchmarking threaten water-sustainability research, and what design principles follow for instrument stewards and universities?
The contribution is threefold. First, the review provides the first systematic comparison of how four families of benchmarking instruments position water (RQ1), synthesised in Table 1. Second, it develops a conceptual framework, presented in Figure 1, that specifies three mechanisms linking benchmarking participation to water research capacity together with the conditions under which each operates (RQ2). Third, it derives risks and design principles that convert the rankings-critique literature into actionable guidance for the water case (RQ3). The remainder of the article reviews the relevant literatures, describes the review method, presents the findings organised around the three questions, and discusses implications, limitations, and the extent to which the identified gap has been closed.
2. Literature Review and Theoretical Background
2.1 Water Sustainability as a Research Imperative
The empirical case for sustained water scholarship is not in dispute. Global assessments document concurrent threats to human water security and to river ecosystems (Vörösmarty et al., 2010), severe scarcity affecting four billion people (Mekonnen & Hoekstra, 2016), and consumption patterns whose water footprint extends far beyond the point of use (Hoekstra & Mekonnen, 2012). Gleick (2003) reframed the policy problem as one of managing demand, improving productivity per unit of water, and matching water quality to use, an agenda that requires interdisciplinary research spanning engineering, economics, and governance. Universities contribute to this agenda in two distinct capacities that the literature tends to conflate: as producers of water research and as water-using organisations. Marinho et al. (2014) illustrate the second capacity, documenting a water conservation programme at a Brazilian public university as a support for wider sustainable practice. The distinction matters for what follows, because benchmarking instruments differ precisely in which of the two capacities they measure. It also has a temporal dimension. Campus water performance can be improved within a budget cycle through metering, retrofits, and behavioural programmes, whereas water research capacity, in the form of laboratories, doctoral pipelines, long-term monitoring sites, and relationships with basin authorities, accumulates over decades and erodes quickly when funding signals turn elsewhere. Any instrument that measures the first capacity while purporting to speak for the second therefore risks rewarding fast, visible improvements at the expense of slow, structural ones. This asymmetry recurs throughout the analysis that follows.
2.2 Global Rankings: Measurement, Validity, and Reactivity
Research on university rankings offers the theoretical core for this review. Espeland and Sauder (2007) established that public measures are reactive: organisations reshape themselves around the measure, so that rankings do not merely describe universities but recreate them. Subsequent scientometric work has specified what the major instruments actually capture. Selten et al. (2020) found that the Academic Ranking of World Universities, THE World University Rankings, and QS rankings are stable over time and that their indicators load principally on two latent factors, institutional reputation and research performance, raising the concern that the variables do not capture the broader quality concepts they claim to measure. Moed (2017) reached a compatible conclusion through a comparative analysis of five world rankings. Vernon et al. (2018), in a systematic review of ranking systems, identified twenty-four systems, evaluated thirteen against basic quality criteria, and reported that no generally accepted indicators exist and that no single system comprehensively evaluates research quality; among systems that disclose weights, the large majority of the weighting rewarded research or teaching quality, often via reputation surveys. Marginson (2014) questioned the social-science validity of ranking constructs, and Gadd (2021) argued that rankings reward institutions that grow up the tables rather than institutions that mature against their own missions. Hicks et al. (2015) distilled the corrective position into principles for responsible metrics, insisting that quantitative indicators should support, not replace, expert judgment. None of this literature, however, asks what these dynamics imply for a specific thematic field such as water. That question requires joining the rankings critique to the sustainability assessment literature.
2.3 Sustainability Assessment in Higher Education
A parallel literature examines how universities institutionalise sustainability and how dedicated instruments assess it. Lozano (2006) analysed the organisational barriers that sustainability initiatives must overcome, and Lozano et al. (2015) found, through a worldwide survey, that declarations of commitment outrun implementation. Findler et al. (2019) reviewed research from 2005 to 2017 and reoriented the field from what universities do toward how their activities affect society, the environment, and the economy. Leal Filho et al. (2019) assessed whether sustainability teaching keeps pace with the SDGs, and Purcell et al. (2019) proposed the university as a living laboratory in which campus operations become sites of research and learning.
Within this field, assessment instruments have received focused attention. Suwartha and Sari (2013) evaluated the early UI GreenMetric ranking as a tool for green university development, and Lauder et al. (2015) subjected the same instrument to a critical review of its design as a global campus sustainability ranking. Alghamdi et al. (2017) compared the indicator sets of sustainability assessment tools across universities, showing substantial variation in coverage. Urbanski and Leal Filho (2015) analysed early institutional data submitted to STARS. Atici et al. (2021) examined empirically how GreenMetric participation relates to standing in world university rankings, connecting the campus-greening and academic-prestige literatures. Galleli et al. (2022) compared UI GreenMetric with the THE World University Rankings against the Berlin Principles and concluded that both exhibit structural differences and methodological limitations, so that institutions must choose instruments contextually rather than expect a single valid measure. For the SDG-specific instruments, De la Poza et al. (2021) used THE Impact Rankings data to model universities' SDG reporting, while Bautista-Puig et al. (2022), analysing the 2019 to 2021 editions, found severe methodological inconsistencies that in their assessment produce a distorted view of sustainability performance, alongside genuine reputational opportunities for less-prominent institutions.
2.4 SDG Measurement and the Research Agenda
A third literature measures how the SDGs are reshaping research itself. Salvia et al. (2019) assessed SDG-related research trends and the alignment between local issues and global agendas. Sianes et al. (2022) traced the scientometric imprint of the SDGs on academic research. Armitage et al. (2020) demonstrated a foundational measurement problem: independent bibliometric approaches to identifying SDG-related publications showed little overlap, and the choice of approach altered country rankings, leading the authors to advise caution toward SDG rankings and tools. At the level of goal politics, Forestier and Kim (2020) documented selective prioritisation among the SDGs by national governments, and Heleta and Bagus (2021) argued that SDG frameworks neglect higher education capacity in low-income countries and can reinforce global inequality.
2.5 Synthesis and Gap
Read together, these literatures supply all the raw material for an account of benchmarking and water research, yet none assembles it. The rankings literature explains reactivity and indicator invalidity but is thematically blind. The sustainability assessment literature evaluates instruments but concentrates on campus operations and aggregate SDG engagement rather than on any single goal's research base. The SDG measurement literature shows that classifying water-related research is itself unstable, which undermines the evidentiary foundation of any water indicator, but it does not follow the consequences into institutional behaviour. What the combined literature has not resolved is (a) a systematic description of where water sits in benchmarking architectures, (b) a specification of the mechanisms by which benchmarking participation could change water research capacity, and (c) an assessment of which known ranking pathologies bear most directly on water. Those three unresolved items correspond to RQ1 through RQ3 and define the contribution of this review.
3. Method
3.1 Review Design
The study uses an integrative review design, which is appropriate when the goal is to synthesise conceptually heterogeneous literatures and to generate new frameworks rather than to estimate pooled effects (Torraco, 2016). Whittemore and Knafl (2005) specify the stages followed here: problem identification, literature search, data evaluation, data analysis, and presentation. The design is complemented by a documentary analysis of the published methodologies of three benchmarking instruments (Times Higher Education, 2024a, 2024b; UI GreenMetric, n.d.; AASHE, n.d.), because indicator architectures are primary documents that the peer-reviewed literature discusses but does not reproduce in full.
3.2 Search Strategy
Searches were conducted in Scopus, Web of Science, and Google Scholar for literature published between January 2003 and August 2026, with the core corpus concentrated in 2013 to 2022. The lower bound admits seminal works on water policy and rankings sociology; the emphasis on the last decade reflects the founding of the THE Impact Rankings and the SDG era. Example search strings, adapted to each database's syntax, included: ("university ranking*" OR "league table*" OR benchmark*) AND (sustainab* OR "SDG*"); ("UI GreenMetric" OR "STARS" OR "Impact Rankings") AND (universit* OR "higher education"); (water OR "SDG 6" OR "water management") AND (universit* OR campus OR "higher education") AND (research OR ranking OR assessment); and ("research assessment" OR bibliometric*) AND ("sustainable development goals"). Reference lists of included articles were snowballed in both directions.
3.3 Inclusion and Exclusion Criteria
Sources were included if they (a) were peer-reviewed journal articles, or methodology documents published by the steward of a benchmarking instrument; (b) addressed at least one of the three review domains, namely rankings and research assessment, sustainability assessment in higher education, or water sustainability and SDG measurement; and (c) reported an identifiable argument, conceptualisation, or empirical result relevant to the research questions. Sources were excluded if they (a) were editorials, theses, or conference abstracts without full analysis; (b) addressed campus sustainability without any assessment, benchmarking, or research-capacity dimension; or (c) could not be verified against their publisher's bibliographic record. Screening proceeded in two stages, title and abstract followed by full text, with the research questions as the screening rubric. Thirty-six sources satisfied all criteria: thirty-two peer-reviewed articles and four instrument methodology documents.
3.4 Analysis and Framework Derivation
Included sources were coded against three analytic categories derived from the research questions: representation (how an instrument defines, weights, and verifies water content), mechanism (any process by which measurement is claimed or shown to change institutional behaviour or research activity), and pathology (any documented distortion attributable to measurement). Constant comparison across the three literatures generated the framework in Figure 1: mechanisms proposed in the rankings and sustainability literatures were retained only where at least two independent sources supported the underlying process, and each mechanism was then specified for the water case using the instrument documents. The framework is therefore a conceptual synthesis, and its water-specific pathways are stated as propositions to be tested, not as established findings.
3.5 Rigour and Trustworthiness
Several safeguards address the known weaknesses of integrative reviews. Selection bias was limited by searching three databases, by snowballing, and by including critical as well as favourable evaluations of every instrument discussed. Verification bias was limited by checking every cited source's bibliographic record against the Crossref or OpenAlex registry and by consulting instrument methodologies in their steward's own publications rather than through secondary description. Interpretive claims are marked as such throughout, and single studies are attributed as single studies. Two scope boundaries should be stated plainly: the review covers documents published in English, which underrepresents scholarship from several water-stressed regions, and it does not attempt a quantitative meta-analysis, because the included studies do not share comparable outcome measures. No systematic-review reporting checklist was applied, and no claim is made about the exhaustiveness of the corpus beyond the stated search protocol.
4. Findings
4.1 Where Water Sits in Benchmarking Architectures (RQ1)
The first finding is structural: across the four families of instruments that dominate global university benchmarking, water occupies positions of sharply different visibility, and in no case is water research capacity measured directly and verifiably. Table 1 summarises the comparison.
Academic world rankings contain no water-specific indicators at all. Their variables reduce, empirically, to reputation and aggregate research performance (Selten et al., 2020), and systematic evaluation finds their indicator sets dominated by research and teaching prestige with no generally accepted standards (Vernon et al., 2018; Moed, 2017). Water research enters these instruments only as an undifferentiated contribution to publication and citation counts. A university could dismantle its entire water institute without any detectable movement in its world ranking position, a property that follows directly from the indicator structure documented in the scientometric literature.
The THE Impact Rankings are the only global instrument with a dedicated water pillar. The SDG 6 methodology allocates 27 percent of the pillar score to research on clean water and sanitation, assessed through Scopus-based citation impact and publication volume for 2018 to 2022, with the remaining weight distributed across water consumption (19 percent), water usage and care (23 percent), water reuse (12 percent), and community engagement (19 percent) (Times Higher Education, 2024b). Two architectural features qualify this apparent prominence. First, participation is selective: universities submit data on as many SDGs as they choose, and the overall score combines SDG 17 with each institution's best three other goals (Times Higher Education, 2024a). SDG 6 therefore competes for attention with sixteen alternatives, and institutions rationally submit their strongest goals, a dynamic of goal-level selectivity that mirrors the cherry-picking Forestier and Kim (2020) documented among national governments. Second, the evidence base is largely self-provided, with missing data scored as zero (Times Higher Education, 2024a, 2024b), and content analysis of the 2019 to 2021 editions found severe methodological inconsistencies that distort the resulting picture of sustainability performance (Bautista-Puig et al., 2022).
The campus-greening instruments treat water as operations. UI GreenMetric includes water usage among its six weighted criteria, alongside setting and infrastructure, energy and climate change, waste, transportation, and education and research (UI GreenMetric, n.d.), on the basis of self-reported campus data whose design has been critically reviewed since the instrument's early years (Lauder et al., 2015; Suwartha & Sari, 2013). STARS awards water-related credits within its operations category in a transparent, self-reporting, points-based framework (AASHE, n.d.; Urbanski & Leal Filho, 2015). In both instruments, water performance means campus water performance; research on water appears, if at all, inside generic education and research credits.
The answer to RQ1 is therefore that benchmarking architectures represent water either not at all, or optionally, or operationally. Only one instrument measures water research, it does so for a self-selected subset of institutions, and it relies on bibliometric classification of water-related publications, a procedure that Armitage et al. (2020) showed to be unstable, since independent SDG mapping approaches produced little overlap and materially different rankings. The validity qualifier in RQ1 is thus answered in the negative: no current instrument offers a valid, comparable measure of water research capacity across institutions.
A comparison across the four architectures also reveals a trade-off between visibility and verifiability. The instruments that make water most visible, the SDG 6 pillar and the campus-greening rankings, rest mainly on self-reported evidence, while the instruments with the most externally verifiable data, the bibliometrics-driven world rankings, make water invisible. The comparative literature reaches a consistent conclusion about this situation: Galleli et al. (2022) found that neither a campus-greening instrument nor an academic ranking satisfies the Berlin Principles fully and advised contextual selection rather than reliance on any single system, and Alghamdi et al. (2017) documented wide variation in which indicators sustainability assessment tools include at all. For water specifically, this means that an institution seeking an external mirror for its water performance must triangulate at least two instruments with different blind spots, and that any single-instrument account of a university's water standing should be treated as partial by construction.
Table 1
Water Sustainability in Four Families of Global University Benchmarking Instruments
Instrument family (steward) | Primary orientation | Placement of water | Water-related indicators and data basis | Key sources |
Academic world rankings (ARWU, THE WUR, QS) | Reputation and aggregate research performance | Absent; no dedicated environmental or water indicators | Water research counted only inside aggregate publication and citation measures; bibliometric data and reputation surveys | Moed (2017); Selten et al. (2020); Vernon et al. (2018) |
THE Impact Rankings (Times Higher Education) | Contribution to the 17 SDGs; overall score combines SDG 17 with each institution's best three other goals | Dedicated, optional SDG 6 pillar | Research on clean water and sanitation (27%); water consumption (19%); water usage and care (23%); water reuse (12%); water in the community (19%); Scopus data plus self-submitted evidence, with missing data scored zero | Times Higher Education (2024a, 2024b); Bautista-Puig et al. (2022); De la Poza et al. (2021) |
UI GreenMetric (Universitas Indonesia) | Campus greening and infrastructure | Water usage as one of six weighted criteria | Campus water usage reported by institutions; weightings under continuous review | UI GreenMetric (n.d.); Lauder et al. (2015); Suwartha and Sari (2013) |
STARS (AASHE) | Institutional sustainability self-assessment | Water credits within the Operations category | Points-based credits including water use; transparent, voluntary self-reporting | AASHE (n.d.); Urbanski and Leal Filho (2015) |
Note. Indicator names, weights, and category placements are taken from the instruments' published methodologies as cited in the final column; characterisations of the academic world rankings reflect the scientometric analyses cited. ARWU = Academic Ranking of World Universities; THE WUR = Times Higher Education World University Rankings; QS = Quacquarelli Symonds; SDG = Sustainable Development Goal; STARS = Sustainability Tracking, Assessment and Rating System; AASHE = Association for the Advancement of Sustainability in Higher Education.
4.2 Mechanisms Linking Benchmarking Participation to Water Research Capacity (RQ2)
The second finding is that, despite these representational weaknesses, the literature supports three distinct mechanisms through which benchmarking participation can plausibly build water research capacity. Figure 1 assembles them into a conceptual framework; each pathway is stated here as a proposition grounded in the sources that support the underlying process.
The first mechanism is signalling and resource mobilisation. Rankings are reactive instruments: organisations reallocate attention and resources toward what is measured (Espeland & Sauder, 2007), and universities have been shown to convert sustainability commitments into structures and budgets unevenly, with visible external commitments outrunning implementation (Lozano et al., 2015). Where an institution elects the SDG 6 pillar, the 27 percent research weighting (Times Higher Education, 2024b) creates, for the first time in any global instrument, a direct reputational return on water scholarship. Bautista-Puig et al. (2022) found that the Impact Rankings offer reputational opportunities precisely to institutions outside the traditional elite, which suggests that the signalling mechanism may operate most strongly for universities in water-stressed middle-income regions, where world rankings offer them little. Proposition 1: universities that elect SDG 6 will, other conditions equal, increase internal allocation to water research relative to observationally similar non-electing institutions.
The second mechanism is measurement and data infrastructure. Benchmarking obliges institutions to meter, audit, and document their own water systems: consumption per capita, wastewater treatment, reuse policies (Times Higher Education, 2024b), and the operational categories of GreenMetric and STARS (UI GreenMetric, n.d.; AASHE, n.d.). This reporting burden creates campus water data that did not previously exist in comparable form, and campus data are the raw material of the living-laboratory model in which operations become research sites (Purcell et al., 2019). The case documented by Marinho et al. (2014), where a university water conservation programme supported wider sustainable practice, illustrates the pathway from operational measurement to applied scholarship. This mechanism reframes the two university capacities distinguished in Section 2.1: benchmarking of the university as water user can subsidise the university as water researcher. Proposition 2: institutions with sustained participation in operations-focused instruments will produce more campus-based water research than non-participants.
The third mechanism is network formation and agenda alignment. SDG-structured benchmarking embeds universities in a shared classification of societal problems, which lowers the cost of identifying partners, and the community components of the SDG 6 pillar explicitly reward off-campus cooperation on water security (Times Higher Education, 2024b). The literature on universities and the SDGs argues that such engagement reorients institutional missions toward societal impact (Findler et al., 2019; Leal Filho et al., 2019), and bibliometric evidence confirms that the SDG framework has left a measurable imprint on research agendas (Sianes et al., 2022; Salvia et al., 2019). Proposition 3: benchmarking participation increases the share of water research conducted with non-academic partners and oriented to local water problems.
The three mechanisms are not independent, and their interactions carry analytical weight. Signalling without measurement produces commitments that outrun implementation, the pattern Lozano et al. (2015) observed across the sustainability declarations of the preceding two decades, because reputational incentives arrive before the data systems needed to act on them. Measurement without signalling produces data that remain administrative, since without reputational or funding stakes there is little pull to convert campus water records into research questions. Network formation amplifies both: partners demand data, which strengthens the measurement pathway, and partnerships generate the demonstrable community engagement that the SDG 6 pillar rewards (Times Higher Education, 2024b), which strengthens signalling. A capacity-building account of benchmarking therefore predicts the strongest effects where all three mechanisms operate together, typically in institutions that elect SDG 6, sustain operational reporting, and hold standing relationships with water authorities.
These mechanisms answer RQ2, but the framework in Figure 1 also specifies their conditions. Each pathway passes through two moderating filters: indicator validity, which is currently weak (Armitage et al., 2020; Galleli et al., 2022), and institutional response type, which ranges from substantive investment to symbolic compliance. The next subsection examines the second filter.

Figure 1. Conceptual framework linking benchmarking participation to water sustainability and management research capacity through three mechanisms (M1 to M3), conditioned by indicator validity and by institutional response. Source: author's elaboration from the reviewed literature.
4.3 Risks: How Benchmarking Can Distort Water Research (RQ3)
The third finding is that every major pathology documented in the rankings literature has a specific and foreseeable water-sector expression.
Reactivity can become gaming. Espeland and Sauder (2007) showed that measured organisations manage the measure, not only the underlying performance. In the water case, the combination of self-provided evidence and zero-scoring of missing data (Times Higher Education, 2024a) rewards documentation capacity as much as water performance, and the inconsistencies identified by Bautista-Puig et al. (2022) indicate that the verification layer is not yet strong enough to separate the two. Institutions with professional rankings offices can therefore outperform institutions with stronger water science but weaker reporting, a concern consistent with the finding of Heleta and Bagus (2021) that SDG frameworks disadvantage under-resourced institutions in low-income countries, including many in the most water-stressed regions.
Classification instability can misdirect credit. The research component of the SDG 6 score depends on bibliometric identification of water-related publications, yet Armitage et al. (2020) found little overlap between independent approaches to exactly this task and showed that the choice of approach changes rankings. Until SDG 6 publication mapping stabilises, research-weighted water scores contain an unquantified layer of classification noise, and universities optimising against them may be optimising against an artefact.
Goal selectivity can hollow out the signal. Because institutions submit their strongest goals (Times Higher Education, 2024a), the population ranked on SDG 6 is self-selected, so pillar positions cannot be read as a census of global water research capacity. Forestier and Kim (2020) showed that selective SDG prioritisation at the national level carries governance costs; the same logic implies that water, a goal requiring expensive infrastructure and specialised research capacity, risks systematic under-election relative to goals that most institutions can document cheaply. This is an interpretive extension of their finding, offered here as a proposition rather than an established result.
Metric fixation can displace judgment. The general corrective is well established: indicators should support expert judgment, not replace it (Hicks et al., 2015), and instruments reward growth up the table rather than maturity against mission (Gadd, 2021). For water, the mission-relevant questions, such as whether research addresses the basin problems of the university's own region, are precisely the ones that Salvia et al. (2019) found imperfectly aligned between local issues and global agendas, and no current indicator captures them.
4.4 Design Principles
Answering the second half of RQ3, four design principles follow from the analysis, each traceable to the evidence above. First, verify before weighting: research-heavy water scores should not exceed the reliability of SDG 6 publication mapping, which argues for published sensitivity analyses across mapping approaches (Armitage et al., 2020). Second, reward disclosure symmetry: instruments should distinguish absent performance from absent documentation, since zero-scoring missing data conflates the two (Times Higher Education, 2024a) and penalises under-resourced institutions (Heleta & Bagus, 2021). Third, connect the operational and research ledgers: instruments already collect campus water data; publishing them in reusable form would let benchmarking directly subsidise living-laboratory research (Purcell et al., 2019; Marinho et al., 2014). Fourth, benchmark contextually: following the conclusion of Galleli et al. (2022) that no single instrument is best, universities should select and interpret instruments against their own water context, and evaluators should follow the principle that metrics inform rather than decide (Hicks et al., 2015).
5. Discussion
5.1 Theoretical Implications
The review's central theoretical claim is that reactivity theory (Espeland & Sauder, 2007) gains explanatory power when it is made goal-specific. Applied at the level of whole institutions, reactivity predicts generic ranking-seeking behaviour. Applied at the level of a single SDG, it predicts differentiated behaviour: election or avoidance of the goal, substantive or symbolic response, and reallocation across goals as relative prices change. The framework in Figure 1 formalises this by treating benchmarking participation as an institutional choice whose research consequences run through three mechanisms and two filters. This specification also connects reactivity theory to the sustainability implementation literature: the gap between declaration and implementation that Lozano et al. (2015) documented is, in the framework's terms, the symbolic branch of the institutional response filter. For the scientometrics of the SDGs, the analysis converts the measurement instability shown by Armitage et al. (2020) from a technical caveat into a theoretical variable, since classification noise determines how much of the signalling mechanism reaches actual water research rather than an artefact of mapping.
5.2 Practical and Policy Implications
For ranking stewards, the analysis implies that the credibility of water pillars now depends less on additional indicators than on verification and on published sensitivity of research scores to mapping choices. For university leaders, the framework offers a decision structure: SDG 6 election is most defensible where the institution can pair reporting with substantive investment, and the measurement mechanism means that even operations-focused participation can be converted into research assets if campus water data are treated as research infrastructure. For funders and water agencies, the self-selected nature of SDG 6 pillar populations means ranking positions should not be used as a screen for research capability; the bibliometric record and expert review remain the appropriate instruments, used under responsible-metrics principles (Hicks et al., 2015). For policymakers in water-stressed regions, the equity findings counsel support for reporting capacity, since otherwise benchmarking will systematically understate the water work of the institutions closest to the problem (Heleta & Bagus, 2021). There is also a positive policy reading of the analysis. Because the measurement mechanism runs through data that universities must collect anyway, national water agencies could treat benchmarked campuses as a distributed observation network: standardised consumption, treatment, and reuse data across hundreds of institutions constitute evidence about demand-side water management of exactly the kind the soft-path agenda requires (Gleick, 2003). Realising that value requires only that instrument stewards publish operational water data in reusable form, a change that costs participants nothing beyond what current reporting already demands.
5.3 Limitations
The limitations follow from the method. First, this is an integrative review: the framework is a conceptual synthesis, its propositions are untested, and no causal claim about benchmarking and water research output is made or warranted. Second, the corpus is English-language and concentrated in journals indexed by the major databases, which underrepresents regions where water stress is most acute and where the equity effects discussed above matter most. Third, the documentary analysis rests on instrument methodologies as published in 2024 editions and current technical manuals; benchmarking methodologies change frequently, and the specific weights cited here will date. Fourth, the review depends in places on single studies, notably for the content analysis of Impact Rankings inconsistencies (Bautista-Puig et al., 2022) and for SDG mapping instability (Armitage et al., 2020); these are attributed as single studies, and the framework would need revision if replication fails. Fifth, no formal quality scoring of included studies was undertaken beyond the stated inclusion criteria, which is a recognised trade-off of integrative designs (Whittemore & Knafl, 2005).
5.4 Future Research
Each proposition in Section 4.2 defines an empirical study. Proposition 1 invites a difference-in-differences design comparing water research investment in SDG 6-electing and non-electing universities, feasible once panel data on pillar participation accumulate. Proposition 2 invites bibliometric analysis of campus-based water research among long-run GreenMetric and STARS participants. Proposition 3 invites co-authorship and funding-acknowledgement analysis of partnered water research before and after benchmarking entry. Beyond the propositions, two measurement studies are prerequisite to all evaluative work: a water-specific replication of the mapping comparison of Armitage et al. (2020) confined to SDG 6 queries, and an audit study of the verification practices behind self-reported water evidence. Finally, qualitative work inside universities is needed to observe the response filter directly, distinguishing substantive from symbolic SDG 6 engagement in the tradition of implementation research (Lozano et al., 2015).
5.5 How Far the Gap Was Closed
Of the three unresolved items identified in Section 2.5, the first, the representational question, is now closed to the extent that public methodologies allow: Table 1 provides the systematic comparison that the literature lacked. The second, the mechanism question, is closed at the conceptual level: the framework specifies pathways and conditions, but their empirical weight remains unmeasured. The third, the pathology question, is closed as translation: known ranking distortions have been given specific water-sector expressions and countermeasures, though several of these expressions are propositions rather than observations. The review therefore converts an unstructured gap into a structured research programme; it does not, and by design cannot, supply the causal evidence that programme calls for.
6. Conclusion
Global university benchmarking has begun to measure water, but it measures it unevenly: not at all in the academic world rankings, optionally and with weak verification in the SDG-based rankings, and operationally in the campus-greening instruments. This integrative review joined the rankings, sustainability assessment, and water research literatures to show what follows from that architecture. Benchmarking participation can strengthen water sustainability and management research through signalling that attaches reputational value to water scholarship, through measurement that creates campus water data usable as research infrastructure, and through networks that align research with local water problems. The same participation can weaken the field through gaming, classification noise, goal selectivity, and metric fixation, and these risks fall hardest on institutions in water-stressed, resource-poor settings. Whether benchmarking advances or distorts water research is therefore not a property of benchmarking as such but of indicator validity and institutional response, the two filters at the centre of the framework proposed here. The propositions derived from that framework set the empirical agenda; the design principles indicate what stewards and universities can change without waiting for it.
Declarations
Funding. This research received no external funding.
Conflicts of Interest. The author declares no conflict of interest.
Ethics. This study is a review of published literature and publicly available documents; it involved no human participants, animals, or personal data, and no ethical approval was required.
Data Availability. No new data were created or analysed in this study. All sources synthesised are cited and publicly available through the references listed.
References
Alghamdi, N., den Heijer, A., & de Jonge, H. (2017). Assessment tools' indicators for sustainability in universities: An analytical overview. International Journal of Sustainability in Higher Education, 18(1), 84–115. https://doi.org/10.1108/IJSHE-04-2015-0071
Armitage, C. S., Lorenz, M., & Mikki, S. (2020). Mapping scholarly publications related to the Sustainable Development Goals: Do independent bibliometric approaches get the same results? Quantitative Science Studies, 1(3), 1092–1108. https://doi.org/10.1162/qss_a_00071
Association for the Advancement of Sustainability in Higher Education. (n.d.). STARS technical manual. Retrieved August 8, 2026, from https://stars.aashe.org/pages/about/technical-manual.html
Atici, K. B., Yasayacak, G., Yildiz, Y., & Ulucan, A. (2021). Green University and academic performance: An empirical study on UI GreenMetric and World University Rankings. Journal of Cleaner Production, 291, 125289. https://doi.org/10.1016/j.jclepro.2020.125289
Bautista-Puig, N., Orduña-Malea, E., & Pérez-Esparrells, C. (2022). Enhancing sustainable development goals or promoting universities? An analysis of the Times Higher Education Impact Rankings. International Journal of Sustainability in Higher Education, 23(8), 211–231. https://doi.org/10.1108/IJSHE-07-2021-0309
De la Poza, E., Merello, P., Barberá, A., & Celani, A. (2021). Universities' reporting on SDGs: Using THE Impact Rankings to model and measure their contribution to sustainability. Sustainability, 13(4), 2038. https://doi.org/10.3390/su13042038
Espeland, W. N., & Sauder, M. (2007). Rankings and reactivity: How public measures recreate social worlds. American Journal of Sociology, 113(1), 1–40. https://doi.org/10.1086/517897
Findler, F., Schönherr, N., Lozano, R., Reider, D., & Martinuzzi, A. (2019). The impacts of higher education institutions on sustainable development: A review and conceptualization. International Journal of Sustainability in Higher Education, 20(1), 23–38. https://doi.org/10.1108/IJSHE-07-2017-0114
Forestier, O., & Kim, R. E. (2020). Cherry-picking the Sustainable Development Goals: Goal prioritization by national governments and implications for global governance. Sustainable Development, 28(5), 1269–1278. https://doi.org/10.1002/sd.2082
Gadd, E. (2021). Mis-measuring our universities: Why global university rankings don't add up. Frontiers in Research Metrics and Analytics, 6, 680023. https://doi.org/10.3389/frma.2021.680023
Galleli, B., Teles, N. E. B., Santos, J. A. R. dos, Freitas-Martins, M. S., & Hourneaux Junior, F. (2022). Sustainability university rankings: A comparative analysis of UI GreenMetric and the Times Higher Education World University Rankings. International Journal of Sustainability in Higher Education, 23(2), 404–425. https://doi.org/10.1108/IJSHE-12-2020-0475
Gleick, P. H. (2003). Global freshwater resources: Soft-path solutions for the 21st century. Science, 302(5650), 1524–1528. https://doi.org/10.1126/science.1089967
Heleta, S., & Bagus, T. (2021). Sustainable development goals and higher education: Leaving many behind. Higher Education, 81(1), 163–177. https://doi.org/10.1007/s10734-020-00573-8
Hicks, D., Wouters, P., Waltman, L., de Rijcke, S., & Rafols, I. (2015). Bibliometrics: The Leiden Manifesto for research metrics. Nature, 520(7548), 429–431. https://doi.org/10.1038/520429a
Hoekstra, A. Y., & Mekonnen, M. M. (2012). The water footprint of humanity. Proceedings of the National Academy of Sciences, 109(9), 3232–3237. https://doi.org/10.1073/pnas.1109936109
Lauder, A., Sari, R. F., Suwartha, N., & Tjahjono, G. (2015). Critical review of a global campus sustainability ranking: GreenMetric. Journal of Cleaner Production, 108, 852–863. https://doi.org/10.1016/j.jclepro.2015.02.080
Leal Filho, W., Shiel, C., Paço, A., Mifsud, M., Ávila, L. V., Brandli, L. L., Molthan-Hill, P., Pace, P., Azeiteiro, U. M., Vargas, V. R., & Caeiro, S. (2019). Sustainable Development Goals and sustainability teaching at universities: Falling behind or getting ahead of the pack? Journal of Cleaner Production, 232, 285–294. https://doi.org/10.1016/j.jclepro.2019.05.309
Lozano, R. (2006). Incorporation and institutionalization of SD into universities: Breaking through barriers to change. Journal of Cleaner Production, 14(9–11), 787–796. https://doi.org/10.1016/j.jclepro.2005.12.010
Lozano, R., Ceulemans, K., Alonso-Almeida, M., Huisingh, D., Lozano, F. J., Waas, T., Lambrechts, W., Lukman, R., & Hugé, J. (2015). A review of commitment and implementation of sustainable development in higher education: Results from a worldwide survey. Journal of Cleaner Production, 108, 1–18. https://doi.org/10.1016/j.jclepro.2014.09.048
Marginson, S. (2014). University rankings and social science. European Journal of Education, 49(1), 45–59. https://doi.org/10.1111/ejed.12061
Marinho, M., Gonçalves, M. do S., & Kiperstok, A. (2014). Water conservation as a tool to support sustainable practices in a Brazilian public university. Journal of Cleaner Production, 62, 98–106. https://doi.org/10.1016/j.jclepro.2013.06.053
Mekonnen, M. M., & Hoekstra, A. Y. (2016). Four billion people facing severe water scarcity. Science Advances, 2(2), e1500323. https://doi.org/10.1126/sciadv.1500323
Moed, H. F. (2017). A critical comparative analysis of five world university rankings. Scientometrics, 110(2), 967–990. https://doi.org/10.1007/s11192-016-2212-y
Purcell, W. M., Henriksen, H., & Spengler, J. D. (2019). Universities as the engine of transformational sustainability toward delivering the sustainable development goals: "Living labs" for sustainability. International Journal of Sustainability in Higher Education, 20(8), 1343–1357. https://doi.org/10.1108/IJSHE-02-2019-0103
Salvia, A. L., Leal Filho, W., Brandli, L. L., & Griebeler, J. S. (2019). Assessing research trends related to Sustainable Development Goals: Local and global issues. Journal of Cleaner Production, 208, 841–849. https://doi.org/10.1016/j.jclepro.2018.09.242
Selten, F., Neylon, C., Huang, C.-K., & Groth, P. (2020). A longitudinal analysis of university rankings. Quantitative Science Studies, 1(3), 1109–1135. https://doi.org/10.1162/qss_a_00052
Sianes, A., Vega-Muñoz, A., Tirado-Valencia, P., & Ariza-Montes, A. (2022). Impact of the Sustainable Development Goals on the academic research agenda: A scientometric analysis. PLOS ONE, 17(3), e0265409. https://doi.org/10.1371/journal.pone.0265409
Suwartha, N., & Sari, R. F. (2013). Evaluating UI GreenMetric as a tool to support green universities development: Assessment of the year 2011 ranking. Journal of Cleaner Production, 61, 46–53. https://doi.org/10.1016/j.jclepro.2013.02.034
Times Higher Education. (2024a). Impact Rankings 2024: Methodology. Retrieved August 8, 2026, from https://www.timeshighereducation.com/world-university-rankings/impact-rankings-2024-methodology
Times Higher Education. (2024b). Impact Rankings 2024: Clean water and sanitation (SDG 6) methodology. Retrieved August 8, 2026, from https://www.timeshighereducation.com/impact-rankings-2024-clean-water-and-sanitation-sdg-6-methodology
Torraco, R. J. (2016). Writing integrative literature reviews: Using the past and present to explore the future. Human Resource Development Review, 15(4), 404–428. https://doi.org/10.1177/1534484316671606
UI GreenMetric. (n.d.). Methodology. Universitas Indonesia. Retrieved August 8, 2026, from https://greenmetric.ui.ac.id/about/methodology
Urbanski, M., & Leal Filho, W. (2015). Measuring sustainability at universities by means of the Sustainability Tracking, Assessment and Rating System (STARS): Early findings from STARS data. Environment, Development and Sustainability, 17(2), 209–220. https://doi.org/10.1007/s10668-014-9564-3
Vernon, M. M., Balas, E. A., & Momani, S. (2018). Are university rankings useful to improve research? A systematic review. PLOS ONE, 13(3), e0193762. https://doi.org/10.1371/journal.pone.0193762
Vörösmarty, C. J., McIntyre, P. B., Gessner, M. O., Dudgeon, D., Prusevich, A., Green, P., Glidden, S., Bunn, S. E., Sullivan, C. A., Reidy Liermann, C., & Davies, P. M. (2010). Global threats to human water security and river biodiversity. Nature, 467(7315), 555–561. https://doi.org/10.1038/nature09440
Whittemore, R., & Knafl, K. (2005). The integrative review: Updated methodology. Journal of Advanced Nursing, 52(5), 546–553. https://doi.org/10.1111/j.1365-2648.2005.03621.x
Hashtags:
.png)






Comments